Overview
Analysis and evolution of observability and cost across large-scale cloud platforms.
Context
In workload-dense environments, observability is both an operational necessity and a material cost driver.
Challenge
Preserve the signals needed for operations and investigation without collecting or retaining data without a clear purpose.
My role
- Log and metric cost analysis
- Kubernetes and CloudWatch troubleshooting
- S3 and Athena alternative evaluation
- Traffic, NAT, transfer and compute sizing analysis
Architecture
Different destinations and retention periods were considered according to operational value, separating frequent queries from occasional investigation.
Technical decisions
- Granularity driven by usage
- Retention proportional to need
- Data-driven analysis before optimization
- Compute, observability and networking evaluated together
Automation
Repeatable queries and processes supported large-consumer identification and alternative comparison.
Engineering challenges
Reducing cost without removing essential signals required understanding access, investigation and operating patterns.
Results
- Clearer cost drivers
- A better visibility-retention balance
- Optimization decisions with technical context