SRE / FinOps / Observability

Cloud Observability & Cost Optimization

Balancing operational visibility, retention, granularity and cost in cloud platforms.

AWSCloudWatchLoggingMetricsKubernetesAmazon EKS
01

Overview

Analysis and evolution of observability and cost across large-scale cloud platforms.

02

Context

In workload-dense environments, observability is both an operational necessity and a material cost driver.

03

Challenge

Preserve the signals needed for operations and investigation without collecting or retaining data without a clear purpose.

04

My role

  • Log and metric cost analysis
  • Kubernetes and CloudWatch troubleshooting
  • S3 and Athena alternative evaluation
  • Traffic, NAT, transfer and compute sizing analysis
05

Architecture

Different destinations and retention periods were considered according to operational value, separating frequent queries from occasional investigation.

06

Technical decisions

  • Granularity driven by usage
  • Retention proportional to need
  • Data-driven analysis before optimization
  • Compute, observability and networking evaluated together
08

Automation

Repeatable queries and processes supported large-consumer identification and alternative comparison.

09

Engineering challenges

Reducing cost without removing essential signals required understanding access, investigation and operating patterns.

10

Results

  • Clearer cost drivers
  • A better visibility-retention balance
  • Optimization decisions with technical context