The clearest sign your cluster autoscaler isn't actually optimizing for cost is a node running at single-digit CPU utilization that's been up for weeks. A cluster-wide snapshot once showed this in aggregate, not just as an anecdote: total pod resource requests were reserving roughly a third of cluster memory, while actual usage sat under a fifth of that. Cluster Autoscaler will scale up when pods can't be scheduled, and it will scale down a node once it's completely empty — but it has no concept of consolidation. If a node is hosting even one small, otherwise-movable pod at the moment the autoscaler checks, that node stays up, fully billed, indefinitely. This is the story of diagnosing that waste precisely, designing a thorough fix for it — and then discovering that the actual fix needed was much cheaper than the thorough one, which is a less common ending than most infrastructure writeups admit to.
Why "scale down empty nodes" isn't bin-packing
Cluster Autoscaler's model is reactive and binary: a node is either needed — it has pods that can't be moved elsewhere without disruption — or it's empty and safe to remove. It never asks the more useful question: could the pods currently spread across four half-empty nodes fit comfortably onto two? That's consolidation — actively repacking nodes that are individually non-empty to reduce the total count — and it sits structurally outside what a reactive scale-down-on-empty controller does by design.
A real snapshot made this concrete:
Node A: [pod 2000m] [pod 500m] ............... 25% utilized
Node B: [pod 8m] .............................. 1% utilized <- this one
Node C: [pod 1800m] [pod 1200m] [pod 900m] .... 40% utilized
Node D: [pod 300m] ............................. 4% utilized
4 nodes billed, roughly 17% average utilization
None of those four nodes is empty, so none of them is a scale-down candidate under the reactive model, even though B and D combined would fit comfortably inside a fraction of a single additional node's capacity.
The root cause goes one layer deeper than "the autoscaler is reactive," though — it's partly a consequence of the default Kubernetes scheduler's own placement policy, which actively spreads new pods across the least-loaded nodes rather than packing them tightly. That's a reasonable default in isolation — it maximizes headroom on any single node — but it systematically fragments a cluster into many partially-full nodes, none of which individually ever crosses the utilization threshold that would make it a scale-down candidate. The scheduler and the autoscaler are each behaving correctly by their own local logic; the waste is an emergent property of the two defaults interacting, not a bug in either one.
The cheapest available fix — tightening pod resource requests so they more accurately reflect real usage — was tried first, and judged explicitly insufficient on its own: high-availability deployments still need enough replicas spread for resilience, the scheduler still spreads by default regardless of how tight the requests are, and a rolling update doesn't "remember" to repack afterward even if it briefly happened to consolidate during the rollout. Tighter requests lower the ceiling on waste. They don't add the one missing mechanism: something that actively, continuously asks whether the current workload could be repacked onto fewer nodes.
What a consolidation-aware autoscaler actually does differently
The alternative autoscaler model doesn't wait for a node to empty out on its own. On a recurring interval, it simulates rescheduling every currently running pod as if the cluster had fewer nodes; if that simulation fits, it picks the cheapest node to vacate, cordons it, evicts its pods — respecting PodDisruptionBudgets and graceful termination windows — lets the scheduler re-place those pods onto the remaining nodes, and only then terminates the now-empty instance. It's a direct, ongoing bin-packing pass against the cluster's current shape, evaluated continuously, rather than a one-time decision frozen at the moment each node was first provisioned.
# Simplified NodePool consolidation policy
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: general
spec:
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 30s
budgets:
- nodes: "10%" # never disrupt more than 10% of nodes at once
- nodes: "0" # freeze disruption during a defined blackout window
schedule: "0 9 * * 1-5"
duration: 1h
limits:
cpu: "1000"
template:
spec:
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand"]
- key: kubernetes.io/arch
operator: In
values: ["amd64"]
WhenEmptyOrUnderutilized is the key phrase: it doesn't just reap nodes that happen to go empty, it actively evaluates nodes that are merely underutilized, below a usage threshold, even while they're still hosting pods. The same six pods from the snapshot above, consolidated:
Node A: [pod 2000m] [pod 500m] [pod 8m] [pod 300m] ......... 85% utilized
Node C: [pod 1800m] [pod 1200m] [pod 900m] .................. 85% utilized
2 nodes billed, roughly 85% average utilization
Same workload, half the nodes — because the system is willing to actively move already-scheduled pods onto a better packing, not just decide whether to add or remove whole nodes at the edges of demand.
Two NodePools, and the parts that must never consolidate
The design went past one generic pool. A CPU-bound parsing workload specifically needed instance families that don't throttle under sustained load — ruling out burstable instance types that are cheaper on paper but degrade badly under a continuous CPU-bound job — and was kept off spot capacity entirely, because that workload's queue-consumer pattern carries a real risk of a stuck, unreclaimed message if an instance disappears with only a couple of minutes' interruption notice. A separate general-purpose pool, running stateless and more disruption-tolerant services, used a wider set of instance families and a much shorter consolidation delay, since evicting and rescheduling those pods is cheap and safe.
Two things were deliberately kept outside the new autoscaler's control entirely, for the same underlying reason: a small, traditionally-managed node group hosts the autoscaler's own controller and core cluster add-ons, because a consolidation controller cannot be allowed to evict itself — if every node in a cluster is under its own management and the controller process dies, there's no mechanism left to launch a replacement node. And the vector database stayed on a traditional, non-consolidated node group too, because its storage volume is zone-pinned and tied to a specific node; a consolidation pass that proactively terminates that node to save cost would force an unplanned storage re-attachment mid-operation, trading a cost optimization for a stateful-service incident. The alternative — annotating specific pods as untouchable so the autoscaler simply skips them wherever they land — was considered and set aside in favor of just keeping that workload off the consolidated pool entirely, which is a simpler invariant to reason about than "untouchable, except we have to remember to keep marking it that way."
A prerequisite surfaced during planning that's easy to miss until it actually bites: the cluster had zero PodDisruptionBudgets configured at the time this was scoped. Without them, an eviction-based consolidation controller can legally take down every replica of a deployment at once, mid-consolidation, with nothing stopping it. Adding PodDisruptionBudgets across the handful of deployments that needed them was treated as a hard prerequisite, not an optional hardening step — the migration plan explicitly could not proceed to turning on aggressive consolidation until that gap was closed.
The planned rollout, deliberately staged rather than big-bang: install the new autoscaler alongside the existing one with zero live workloads pointed at it, migrate the most resilient worker pool first — one whose queue-consumer pattern already tolerates message redelivery safely, so an unexpected eviction mid-job is a non-event rather than an incident — then migrate general stateless workloads incrementally, and only decommission the old autoscaler once every pool that could safely move had moved. The stateful, zone-pinned workload was never scheduled to move at all.
The plot twist: the migration got planned, and then deliberately not executed
Here's the part that most writeups about this kind of project leave out, because it doesn't flatter the thing being written about: this migration was fully designed — the NodePools, the disruption budgets, the staged rollout plan, the projected cost savings — and then it was not executed, repeatedly, for a specific and defensible reason.
Projected numbers at the time looked compelling on paper: a baseline cost after an interim right-sizing pass, a conservative estimate for consolidation alone with no spot usage at roughly a 29% reduction, and a more aggressive estimate mixing in spot capacity for safely-interruptible workloads landing around 31%. A real migration project at that scope was estimated at four to six weeks of focused engineering time.
What actually happened instead: a scripted manual runbook. Cordon one node, drain it, and let the existing, completely unmodified autoscaler clean up the now-empty node on its own within five to ten minutes — because the old autoscaler was never bad at removing empty nodes, it was only ever bad at proactively creating emptiness in the first place. The manual drain supplies exactly the one missing step; everything downstream of that was already working.
# The entire "migration" that actually shipped, one node at a time:
kubectl cordon node-i-0123456789abcdef0
kubectl drain node-i-0123456789abcdef0 --ignore-daemonsets --delete-emptydir-data
# wait for pods to reschedule, then let Cluster Autoscaler notice the
# now-empty node and remove it on its own — no custom controller needed
The runbook that made this safe to repeat carries three explicit warnings, each one a documented near-miss: never drain more than one node at a time — draining in parallel causes a burst of pods to go Pending simultaneously, which makes the autoscaler spin up new nodes to absorb them, a net loss; the zone-pinned stateful workload's node gets drained last, or skipped entirely, never in the middle of a batch; and never delete a node or terminate its instance directly — the underlying autoscaling group would simply launch a replacement to maintain its configured desired count, since only the autoscaler itself is actually allowed to lower that count. kubectl drain first, always, never a shortcut straight to termination.
After a few rounds of this manual process, combined with other unrelated cost-reduction work, the general-purpose pool's node count came down to as few as two. At that point the tracked status was blunt about it: node count is low enough without the new autoscaler that a four-to-six-week migration isn't justified yet — revisit if node count creeps back up. Karpenter adoption became explicitly conditional on a measured regression signal, checked on a recurring cadence, rather than something committed to on a fixed calendar date regardless of whether it was still needed.
The actual lesson, and it isn't "use Karpenter"
It's tempting to write this up as "migrate your autoscaler, cut costs by a third." The more honest and more broadly useful lesson is the discipline underneath it: diagnose the real mechanism causing the waste — in this case, a scheduler that spreads by design and an autoscaler that only removes nodes that are already empty, interacting to produce fragmentation neither one intended — design the thorough, correct fix for it in enough detail to know its real cost and its real projected benefit, and then check whether a much cheaper, mostly-manual intervention gets you close enough to the same outcome before committing engineering weeks to the thorough version. Sometimes it does. When it does, the right call is to take the cheap win, write down exactly the signal that would justify revisiting the expensive one, and actually watch for that signal — not to build the sophisticated system just because you'd already finished designing it.

