Autoscaling has a reputation as the thing that saves you money, and that reputation is half-earned. It saves money when it’s scaling down workloads that were genuinely idle. Just as often it papers over the real problem, which is that the thing was sized wrong to begin with, and now you’re paying to automatically provision too much of it on demand. Autoscaling a badly-sized service just means you’re wrong at variable scale.
Most of the money in a spiky workload is made before autoscaling ever runs, in the unglamorous work of figuring out what one unit of your service actually needs. Get the size right first. Then let scaling handle the bursts.
Right-sizing comes before autoscaling, always
The order matters and people get it backwards. If you turn on autoscaling before you know your true per-task footprint, you scale the wrong number. Too generous and every scale-out event over-provisions and the floor costs more than it should. Too tight and it thrashes, adding and removing tasks constantly, which is its own kind of expensive and unstable.
So first, watch a single task under real traffic for long enough to see the actual pattern:
- What does it use at genuine idle? That’s your floor, and it’s usually lower
than the round number someone picked at setup.
- What does it use serving a real request? That’s your per-request cost.
- Where’s the real constraint: CPU, memory, or connections? You size to the one
that binds first, not to all three equally.
Almost every account I’ve looked at has services provisioned to a comfortable round number that bears no relationship to what they use. A task sitting at three percent CPU on a generous allocation is not “safe headroom,” it’s a standing charge for nothing. Shrink it to what it actually needs plus honest margin, and you’ve cut the cost of every single copy of it, at every scale.
Then scale on the metric that actually binds
Once a task is sized honestly, scale on the signal that reflects real pressure. The default reflex is average CPU, and for plenty of services that’s fine. But if your service is bound by something else, CPU-based scaling will sit there looking calm while the thing that actually matters saturates.
Target tracking on the binding metric:
- CPU-bound service -> target ~60% average CPU
- request-bound service -> target requests-per-target on the ALB
- latency-sensitive -> scale before p95 degrades, not after
The number to internalise: scale on the leading indicator, not the lagging one. If you wait until latency has already degraded to add capacity, your users felt the problem before your infrastructure reacted. Target a utilisation that leaves room for the cold-start delay of new tasks, because new capacity is never instant.
Mind the floor and the cold start
Two settings decide most of the cost and most of the risk:
The minimum. Your floor is what you pay around the clock, so it should be the smallest count that serves your genuine baseline. But zero, or too low, and your first burst hits before capacity arrives. For unpredictable traffic, set the floor to your real quiet-hours demand and no higher.
The cold start. New tasks take time to pull an image, boot, and pass health checks. During that window the crowd is already growing. This is the single biggest reason autoscaling feels like it “didn’t work”: it did work, it just couldn’t beat physics. Two honest responses. For unpredictable spikes, keep enough warm headroom that scaling has time to catch up. For predictable ones, raise the floor before the event and lower it after, which is cheaper and more reliable than hoping the scaler wins the race.
Non-production doesn’t need to run all night
The easiest money on the whole bill: development and staging environments running twenty-four hours a day for people who work eight. A scheduled scale-to- zero overnight and on weekends takes a large bite out of non-production cost with essentially no risk, because nobody’s using them at 3am anyway.
Scheduled scaling on dev/staging:
- weekday 07:00 -> min 1
- weekday 19:00 -> min 0
- weekends -> min 0
It’s almost embarrassingly simple, and it’s often the biggest single line-item win available, precisely because it’s so boring nobody bothered.
The boring conclusion
Autoscaling is a burst-handling tool, not a substitute for knowing your service. Right-size a single task against real usage first, so every copy you ever run is sized honestly. Scale on the metric that actually binds, not the one that’s easiest to read. Set the floor to your true baseline, respect the cold start instead of pretending it isn’t there, and switch off the environments nobody’s using at night. None of it is clever. All of it moves the number, which is rather the point.


Leave a Reply