Why Karpenter changes the economics
Cluster Autoscaler scales groups of identical nodes. Karpenter provisions individual nodes, chosen just in time from the whole EC2 catalogue to fit the pods that are pending, and then keeps looking for cheaper ways to run what is already scheduled. That second part, consolidation, is where most of the savings live. It only works if you let it.
Three principles drive every recommendation below:
- Constrain less. The wider the set of instances Karpenter may choose, the cheaper and more Spot-resilient the fleet.
- Disrupt safely, not rarely. Consolidation requires moving pods. Make that safe with PodDisruptionBudgets instead of turning it off.
- Separate what you commit to from what you burst on. Weighted NodePools let a Savings Plan cover the floor while Spot covers the rest.
NodePool requirements: widen the funnel
The most common anti-pattern is a NodePool that whitelists two or three instance types, which is exactly how the old node group behaved. Karpenter cannot save money it is not allowed to find. Express requirements as categories, generations and sizes, not as a list of types, and allow both Spot and On-Demand and both CPU architectures.
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: general
spec:
template:
spec:
nodeClassRef:
group: karpenter.k8s.aws
kind: EC2NodeClass
name: default
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"]
- key: kubernetes.io/arch
operator: In
values: ["arm64", "amd64"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["c", "m", "r"]
- key: karpenter.k8s.aws/instance-generation
operator: Gt
values: ["4"]
- key: karpenter.k8s.aws/instance-size
operator: NotIn
values: ["nano", "micro", "small", "metal"]
expireAfter: 720h # rotate nodes monthly for fresh AMIs
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 1m
budgets:
- nodes: "10%" # never disrupt more than 10% of nodes at once
- nodes: "0" # freeze disruption during business-critical hours
schedule: "0 9 * * mon-fri"
duration: 8h
limits:
cpu: "2000"
memory: 4000Gi
Notes on the choices above:
- Generation > 4 excludes old, expensive-per-unit-of-work families while staying open to everything new.
- Excluding tiny sizes avoids nodes where DaemonSets eat most of the capacity. Excluding
metalavoids surprise giants. - Both architectures require multi-arch images. Build them once; Graviton is one of the biggest levers you have.
- Limits are a safety net against runaway scheduling, not a sizing tool. Set them well above normal peak.
Consolidation: the setting that pays the bill
WhenEmptyOrUnderutilized tells Karpenter to remove empty nodes and to replace or remove underutilised ones when the pods could fit elsewhere more cheaply. This includes Spot-to-Spot consolidation, where Karpenter replaces a Spot node with a cheaper Spot node when it has enough instance-type diversity to do so safely.
Teams that see “no savings” from Karpenter have almost always done one of three things:
- Set
consolidationPolicy: WhenEmpty, so nodes are only removed when nothing runs on them, which on a busy cluster is never. - Shipped workloads without PodDisruptionBudgets, or with PDBs of
maxUnavailable: 0, making every node undrainable. - Annotated far too many pods with
karpenter.sh/do-not-disrupt: "true". Reserve that for genuinely non-interruptible work such as a long batch step.
A PodDisruptionBudget is not an availability feature you add later. In a Karpenter cluster it is the mechanism that makes savings possible.
Disruption budgets give you the safety to run consolidation continuously. Cap concurrent disruptions to a percentage of nodes, and schedule a zero-disruption window during peak traffic or deployment freezes rather than switching consolidation off cluster-wide.
Spot done properly
- Diversify. Karpenter uses the price-capacity-optimized allocation strategy, which works best when it has dozens of candidate types. That is another reason to specify categories, not types.
- Handle interruptions natively. Enable Karpenter’s interruption handling by giving it an SQS queue fed by EventBridge for Spot interruption, rebalance and health events. It will cordon, drain and replace the node before EC2 reclaims it.
- Keep stateful and singleton workloads on On-Demand using a
nodeSelectoronkarpenter.sh/capacity-type: on-demand, and let everything else float. - Spread replicas with
topologySpreadConstraintsusingScheduleAnywayacross zones and, if useful, across capacity types, so a Spot reclaim never takes every replica at once.
Weighted NodePools: commitments and bursts, side by side
You will usually still want a Compute Savings Plan for the steady baseline. The trick is to make Karpenter fill that committed capacity first and burst onto Spot only above it. Two NodePools with weights and limits do exactly this:
# 1) On-Demand baseline, sized to the Savings Plan you own
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: baseline-od
spec:
weight: 100 # preferred while under its limit
limits:
cpu: "400" # ≈ the vCPU your Savings Plan covers
template:
spec:
nodeClassRef: { group: karpenter.k8s.aws, kind: EC2NodeClass, name: default }
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["c", "m", "r"]
- key: karpenter.k8s.aws/instance-generation
operator: Gt
values: ["4"]
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 1m
---
# 2) Spot overflow with no limit and lower weight
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: burst-spot
spec:
weight: 10
template:
spec:
nodeClassRef: { group: karpenter.k8s.aws, kind: EC2NodeClass, name: default }
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["spot"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["c", "m", "r"]
- key: karpenter.k8s.aws/instance-generation
operator: Gt
values: ["4"]
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 1m
Karpenter prefers the higher-weight pool until its limit is reached, then spills into Spot. Measure the baseline honestly before you buy the plan; see Closing the Request Gap for why the baseline should shrink first.
EC2NodeClass: the details that leak money
- Root volumes: set
blockDeviceMappingsexplicitly to gp3 with a size that matches image footprint. Defaults are often larger and slower per dollar than needed. - AMI selection: pin to an alias such as
al2023@latestor a Bottlerocket alias and letexpireAfterplus drift rotate nodes. Patch management becomes free. - Subnets and security groups: use discovery tags so the fleet can land in every private subnet and zone. Fewer zones means fewer Spot pools.
- Instance metadata and tags: propagate cost-allocation tags so the nodes Karpenter creates are attributed correctly in the Cost and Usage Report.
Anti-patterns we keep finding
| What we see | Why it hurts | Fix |
|---|---|---|
| Instance-type whitelists | Blocks cheaper types and starves Spot diversity | Use category, generation and size requirements |
WhenEmpty consolidation | Nodes are never empty on a busy cluster | WhenEmptyOrUnderutilized with budgets |
No PDBs or maxUnavailable: 0 | Nothing can be drained, so nothing consolidates | PDBs with realistic minAvailable |
| Inflated pod requests | Karpenter buys exactly what pods ask for | Right-size requests first |
| One NodePool for everything | Cannot align commitments or isolate GPU and stateful work | Weighted, purpose-specific NodePools |
| Single-zone subnets | Fewer Spot pools, more interruptions, more On-Demand fallback | Discover subnets in every zone |
Very short consolidateAfter on bursty clusters | Node churn and cold starts outweigh savings | Tune to minutes, watch churn metrics |
Measure it
Karpenter exposes Prometheus metrics for provisioning decisions, disruption actions and, in recent versions, per-NodePool cost estimates. Track:
- Spot share of vCPU hours and how often Spot falls back to On-Demand.
- Consolidation actions per hour, and pending-pod time, to balance savings against churn.
- Node utilisation (allocatable versus requested versus used), ideally through Kubecost or OpenCost alongside the Karpenter metrics.
- Baseline pool utilisation against the Savings Plan commitment, so you know when to resize the plan or the limit.
Karpenter cost checklist
- Requirements by category, generation and size, not by type.
- Spot and On-Demand, arm64 and amd64, all zones.
WhenEmptyOrUnderutilizedwith disruption budgets, notWhenEmpty.- PodDisruptionBudgets on every Deployment;
do-not-disruptonly where justified. - Native interruption handling via SQS and EventBridge.
- Weighted On-Demand baseline pool sized to your Savings Plan, Spot overflow pool unlimited.
- gp3 root volumes, AMI alias plus
expireAfterfor rotation. - Pod requests right-sized before you judge the results.
Configured this way, Karpenter stops being a smarter autoscaler and becomes a continuous procurement engine that renegotiates your compute every minute. That is a very different cost profile from the node groups it replaced.