Use values from the alert and earlier steps
Actions accept templates. An unresolved reference fails the step with an error and never sends a blank value.
One runbook can serve many near-identical workloads this way. A step that names
{{alert.labels.service}} restarts whichever service alerted. An alert from a synthetic check linked to a service carries that service as a label, and an SLO breach carries sli_service, the service its SLI is tagged with, so choose that value deliberately.
Fan out over many items
A discovery step (get_unhealthy_pods, get_unhealthy_deployments, or an AWS list_* action) finds items, and a later write step can act on each one. Tick Fan out: run this action once per item found by the discovery step and set the identifying field to *. If discovery finds nothing, the step does nothing. If discovery finds many items, the step pauses and asks for approval with the count, in the A person at each risky step mode. There is also a limit on how many items one fan-out may touch, and above it the step refuses instead of asking, in every mode. After the fan-out can re-scan once the writes finish and fail the step unless fewer, or none, remain, with a wait of up to 600 seconds for changes AWS applies in the background.
Verify what a step says
A step passes when its connector answers without an error. To fail on what the answer says, set Verify Output on the step: pick an operator and, for most, a value. The operators arecontains, not contains, equals, not equals, regex, items empty, items present, count eq and count lt. The item and count operators read what the step found rather than what it printed, so “the re-scan must find nothing” is items empty with no text to type.
Verify with the same signal that triggered the runbook. If a synthetic check started it, finish with a step that runs that check and expects success.
What each connector can do
Steps run through the connectors you configure under Settings > Infrastructure. The four AWS connector types share one action set.prometheus, cloudwatch and synthetic use your existing data sources and checks, so they need no connector. Actions that run arbitrary commands (exec_command, run_command, ssm_send_command) should be marked high risk.
Two actions look alike and are not. restart_deployment restarts one named deployment and refuses an empty name. restart_unhealthy_deployments restarts every unhealthy deployment in the namespace, or in the whole cluster if you give no namespace.
When a step fails
Rules of thumb
- If the platform cannot do a step, use an
inputorapproval_gatestep and hand off to a person. - Put the destructive step behind exactly one approval, and restart the traffic-facing component last, after the services behind it.
- Prefer a recipe over hand-written steps whenever one fits.
Related
- Connect AWS with a read-only role: give a runbook connector read access.
- Monitor endpoints with synthetic checks: verify a step with a synthetic check.
- Troubleshooting: more refused-step messages and fixes.

