Autoscaling on CPU: minimum, maximum, target

Reading · 7 min · Module 7, lesson 3 of 537 min left in this module

Module 7 · ScalingLesson 3 of 5

Goal: Set a service's autoscaling minimum, maximum and CPU target so it grows with load and stays inside what you can pay for.

3:46 · captions and chapters · narrated with an AI-generated voice
Transcript

Narration uses an AI-generated voice.

[00:00] Where we're going

By the end of this video, you'll set three numbers, and watch a service grow from one spherelet to two under load. Then shrink back, once the load is gone.

[00:11] Three numbers

Autoscaling changes the spherelet count for you, using three numbers. First, the target. That's the average CPU you want each spherelet to run at. The default is eighty percent. Second, the minimum. That's the count at quiet times. Third, the maximum. That's the ceiling, and so the most it can cost. The shape stays the same. A Flex service gets more Flex spherelets, never a bigger one.

[00:40] How it decides

Here's how it decides. Say the target is fifty, and one spherelet is working flat out. That's well above fifty, so it adds a spherelet. The work stays on the first one, and the new one is idle. One hot, one idle averages about fifty. Right on target. So it holds at two, which is also the maximum. When the load ends, the average drops. Once it stays low, it removes spherelets, down to the minimum.

[01:10] Choosing them

So how do you choose? A new spherelet takes time to start and pass its health check. Until then, the others carry the load. A target of fifty to seventy percent leaves them room for that wait. At ninety, they're saturated before help arrives. For the minimum, one is the cheapest. Two keeps the service up if one spherelet fails. And set the maximum from what you can afford. Every spherelet counts against your account, and is billed.

[01:41] Turn it on

Let's turn it on. Open the service, then Settings. On the Spherelets card, choose Edit. Turn on Enable Auto-Scaling, and three fields appear. The CPU Utilization Threshold is the target. We'll use fifty. Min Spherelets, one. Max Spherelets, two. The slider goes up to thirty six. It doesn't stop at what you have free, so work that out yourself.

Choose Save. Changes apply on the next deploy, so choose Save and redeploy. The card now shows Auto Scaling Enabled, a CPU target of fifty percent, and a range of one to two. Autoscaling comes with the Team plan and above.

[02:27] Watch it scale

Now for some load. We kept one spherelet busy, burning CPU for five minutes. On the Metrics tab, pick the last fifteen minutes, and Apply. The Spherelets chart steps from one to two. Adding is quick. Overview shows it too. Two of two healthy.

Then the load ended. Removing is deliberately slow. The load has to stay low first. Here, it dropped back to one about six minutes later. So a short lull doesn't throw away spherelets you're about to need.

[03:02] What it won't do

One last thing. Autoscaling looks at CPU only. It won't add spherelets for a memory-hungry app. And if requests are slow because a database is slow, more copies just send it more queries.

[03:17] Recap and your lab

So, a quick recap. The target is the average CPU per spherelet. Fifty to seventy leaves room. The minimum is the quiet count. The maximum is your budget. And adding is quick, while removing is slow, on purpose. Now it's your turn. In the lab, you'll scale the shop out by hand, then let autoscaling take it from one to two and back. I'll see you there.

Key idea

Autoscaling changes the spherelet count for you. When the average CPU across a service's spherelets stays above your target, ComputeSphere adds spherelets, up to your maximum; when it falls, it removes them, down to your minimum.

The three numbers

  • Target (the console's CPU utilization threshold): the average CPU you want each spherelet to run at, as a percentage of its shape's vCPU. The default is 80.
  • Minimum (Min spherelets): the count at quiet times. It never goes lower.
  • Maximum (Max spherelets): the ceiling, and so the most it can cost.

The service keeps its shape. Autoscaling adds more Flex spherelets to a Flex service; it never moves it to Standard.

Choosing them

Target. A new spherelet takes time to start and pass its health check, and the others carry the load until it does. A target of 50 to 70% leaves them room to absorb that wait. At 90%, they're already saturated before help arrives. On the Metrics tab, the Avg view of the CPU chart (titled CPU · avg) shows the per-spherelet figure the autoscaler works from.

Minimum. 1 is the cheapest. 2 keeps the service up if one spherelet fails, which the next module covers.

Maximum. Set it from what you can afford, not from hope. Every spherelet counts against your account's capacity and is billed. Don't count on the console's slider to stop at the spherelets you have free: work that number out yourself, and set the maximum at or below it. On a trial account, with two Flex spherelets, that's 2, and only if nothing else is using them.

How it behaves

Adding is quick: on a test run, a second spherelet came about a minute after CPU crossed the target. Removing is deliberately slow: the load has to stay low first, and the same run dropped back about six minutes after the load ended, so a short lull doesn't throw away spherelets you're about to need.

Autoscaling looks at CPU only. It won't add spherelets for a memory-hungry app, and it won't help when requests are slow because a database is slow: more copies just send it more queries.

Turning it on

Open the service, then Settings. On the Spherelets card choose Edit, turn on Enable Auto-Scaling, and set the three numbers. Choose Save, then Save & redeploy: spherelet and autoscaling changes apply on the next deploy, and Save only waits for one. The card then shows Auto scaling Enabled, the CPU target, and the range, such as 1 — 2, under a second Spherelets label.

Why not csph or computesphere.yaml?

csph has no autoscaling settings: csph deployments autoscale is another name for scale, which sets a fixed count. A manifest's autoscaling_enabled, autoscaling_min, autoscaling_max and autoscaling_cpu_target fields apply when a service is created, but re-applying them doesn't change an existing service yet. Use the console to change them.

Check yourself

A service autoscales between 1 and 4 spherelets with a target of 60%. It runs 2 spherelets, each at about 90% CPU. What happens next?
Your API is slow because every request waits 2 seconds on a database. CPU is at 15%. Will autoscaling help?