Key idea
An alert rule watches one environment. It fires when any spherelet in it stays above a CPU or memory threshold for the whole evaluation period, and it notifies you only while the project's alerts are switched on.
Where rules live
Open the project, choose the ⋯ menu next to its name, then Alerts. The page lists Alert rules; choose New alert to add one.
A rule belongs to an environment, not a service. It watches every spherelet of every service in that environment, each measured against its own shape.
The fields
- Alert type: CPU or Memory.
- Threshold (%): a percentage of the spherelet's shape. On Flex (0.25 vCPU, 512 MB), CPU 80 means 0.2 vCPU; Memory 90 means about 460 MB.
- Evaluation period (seconds): how long it must stay above the threshold before the alert fires. It's counted in whole minutes, with a minimum of two: 300 means five minutes, and anything under 180 means two.
- Severity: Low, Medium or High. It labels the alert; choose it by how fast someone must act.
- Environment: the one to watch.
Choose Create. The rule appears with its own on/off switch: turn it off to pause the rule and keep it for later, or use the bin icon to delete it.
Switch the project's alerts on
A rule notifies anyone only while the project's alerting is on. Once a rule exists, an Alerts on switch appears in the Alert rules header. Turn it on and confirm with Turn on alerts.
From the terminal, csph alerts create makes the same rule (csph context shows the IDs of your default project and environment):
csph alerts create --project <project-id> --environment <environment-id> \
--alert-type cpu --severity medium --threshold 80 --evaluation-period 300
The project's Alerts on switch is in the console only.
Settings that won't flap
- CPU: 80% for 300 seconds. A burst of traffic passes; a service that's been starved for five minutes is worth a look.
- Memory: 90% for 300 seconds. Memory near the ceiling for that long usually ends in Out of memory (lesson 6.3.2).
If a rule fires and nobody needs to act, raise the threshold or the period rather than learning to ignore it.
How CPU is measured, and deleting a project
For CPU, usage is averaged over the evaluation period and compared with the threshold, so a short spike gets smoothed out twice.
Delete a project's alert rules before you delete the project: rules aren't removed with it.
Check yourself