The request gap is the whole game
Kubernetes schedules on requests. The autoscaler adds nodes when requested CPU and memory no longer fit. Your cloud bill, therefore, is a function of the sum of requests across the cluster, not the sum of actual usage. In a typical cluster we audit, actual utilisation sits far below what has been requested. That gap is paid for in full, every hour, and it is invisible in Cost Explorer.
Three layers stack on top of each other:
- Usage → Request gap: pods request more than they use.
- Request → Node gap: nodes are not packed tightly, leaving stranded capacity.
- Node → Price gap: the nodes you do run are bought at the wrong price (On-Demand, wrong family, no commitments).
Optimise in that order. Cheaper nodes that are still mostly empty are still waste.
Layer 1: right-size requests
Measure before you cut
Run the Vertical Pod Autoscaler in recommendation-only mode (updateMode: "Off") across every namespace, and compare its targets with the current requests. Tools such as Goldilocks give the same view with a dashboard. Look for the classic patterns:
- Requests copied from a Helm chart default and never revisited.
- CPU requests set to the same value as limits “to be safe”.
- Java services requesting memory for a heap that was tuned down years ago.
- Sidecars (service mesh, log shippers) requesting as much as the application.
Set requests to P95 usage, plus headroom, and drop CPU limits
Memory requests should sit near real peak usage with a safety margin, because memory is not compressible and OOM kills are outages. CPU is compressible: set requests to a realistic P95 and, for most services, remove CPU limits entirely. CPU limits cause throttling on burst without saving a cent, since the node is billed regardless.
resources:
requests:
cpu: "250m" # P95 of observed usage
memory: "512Mi" # peak usage + ~20% headroom
limits:
memory: "512Mi" # memory limit == request keeps QoS predictable
# no CPU limit: let bursts use idle cycles on the node
Enforce defaults with LimitRange and ResourceQuota
Every namespace needs a LimitRange so pods without requests get sane defaults, and a ResourceQuota so a single team cannot silently double the cluster. Requests are a budget; treat them like one.
Layer 2: pack the nodes
Give the scheduler room to consolidate
- PodDisruptionBudgets on every Deployment. Without them, autoscalers will not evict pods to drain a half-empty node, and consolidation never happens.
- Avoid unnecessary anti-affinity and
topologySpreadConstraintswithDoNotSchedule. PreferScheduleAnywayunless the availability requirement is real. - Watch DaemonSet overhead. Ten DaemonSets on a small node can consume a third of it before any workload lands. Bigger nodes dilute that tax; fewer DaemonSets remove it.
- Reduce pod count per service where fewer, larger replicas serve the same load. Every replica carries fixed sidecar and runtime overhead.
Let the autoscaler choose the node
Static managed node groups with one instance type are the enemy of bin packing. Replace them with Karpenter, which provisions the cheapest node that fits the pending pods and actively consolidates underutilised ones. If you must keep Cluster Autoscaler, use multiple node groups with mixed instance types and enable the least-waste expander.
Layer 3: buy compute well
| Lever | Applies to | Notes |
|---|---|---|
| Spot | Stateless, replicated, batch | Diversify across many instance types and AZs; handle interruptions with the Node Termination Handler or Karpenter’s native handling. |
| Graviton (arm64) | Almost everything with a multi-arch image | Better price-performance than x86 equivalents. Build multi-arch images once; mix architectures in the same cluster. |
| Savings Plans | The measured On-Demand baseline | Compute Savings Plans follow the workload across families and regions. Cover the floor, never the peak. |
| Newer generations | Everything | Each generation is typically cheaper per unit of work. Stop pinning to m5 out of habit. |
Order matters: right-size first, pack second, commit last. A Savings Plan bought against an inflated baseline locks the waste in for three years.
Non-production is where the easy money is
- Scale dev and staging to zero outside working hours. KEDA cron scalers or a simple scheduled job that sets replicas to zero, plus a node autoscaler that removes empty nodes, removes a large share of non-production spend overnight.
- Ephemeral preview environments per pull request that are destroyed on merge beat long-lived “QA2” clusters every time.
- One cluster, many namespaces for non-production. The per-cluster control plane fee is small, but the duplicated DaemonSets, ingress controllers and observability stacks are not.
- Spot everywhere in non-production. Interruptions there are free chaos engineering.
The network and storage lines nobody looks at
- NAT Gateway data processing. Image pulls, S3 reads and package downloads through NAT add up fast. Add VPC endpoints for S3, ECR and the AWS APIs you use, and use ECR pull-through caches for public images.
- Cross-AZ traffic. Chatty services spread across three AZs pay for every byte between them. Enable Topology Aware Routing so Services prefer same-zone endpoints, and colocate tightly coupled workloads.
- Load balancers. One ALB per Ingress is the default and it is expensive. Use ingress grouping with the AWS Load Balancer Controller to share ALBs.
- EBS volumes. Migrate gp2 to gp3, shrink oversized PVCs, and find orphaned volumes from deleted StatefulSets. Set a reclaim policy you actually mean.
- Observability. High-cardinality metrics and debug logs in CloudWatch or a SaaS vendor can rival compute spend. Sample, aggregate and drop labels you never query.
Autoscaling the workloads themselves
Node autoscaling only removes nodes that are empty. Workload autoscaling is what empties them.
- HPA on CPU or memory for request-driven services, with sensible
minReplicas. Three replicas for availability, not ten for comfort. - KEDA for event-driven workloads: scale consumers on queue depth and, crucially, scale to zero when the queue is empty.
- Jobs and CronJobs with
ttlSecondsAfterFinishedso finished pods do not hold requests.
Make cost visible where engineers work
Cost Explorer shows instances, not teams. Install Kubecost or OpenCost, or enable split cost allocation data for EKS in the Cost and Usage Report, so spend is attributed to namespaces and labels. Then put a cost panel next to the latency panel on every team dashboard. Engineers optimise what they can see.
First-month checklist
- Deploy VPA in recommendation mode; fix the ten worst request offenders.
- Remove CPU limits; set memory request equal to limit.
- Add PodDisruptionBudgets everywhere; relax hard anti-affinity.
- Replace static node groups with Karpenter; enable consolidation.
- Build multi-arch images and move to Graviton where possible.
- Spot for all non-production and for stateless production.
- Scale non-production to zero at night; scale queue consumers to zero when idle.
- Add VPC endpoints, share ALBs, enable Topology Aware Routing.
- Install Kubecost or OpenCost; show cost per namespace to the teams.
- Only then: buy a Compute Savings Plan for the new, smaller baseline.
Kubernetes does not make cloud expensive. Unmanaged requests do. Close the request gap and the rest of the bill follows.