Alerts are not arriving
| Symptom | Cause | Fix |
|---|---|---|
The sender reports 404 and “Invalid webhook token”. | The URL has the wrong token, the token was rotated and the old URL has run out, or your organization is disabled. | On Integrations, Webhooks tab, copy the current URL into the sender. After Rotate token, the old URL works for 72 hours. |
The sender reports 401. | A signing secret is saved for the source and the request is unsigned or signed with a different secret, or signing is required for your organization. | For PagerDuty or GitHub, make the secret in the sender match the one on the Webhooks tab, or save the sender’s secret there again. For other sources, contact support. |
The sender reports 429. | You are over the request rate for the endpoint, or one address sent many failed requests in a minute. | Slow the sender down and retry after the retry-after time. Fix the URL if requests are failing authentication. |
| A request is refused with HTTP 403 before it reaches SRE Agent. | The request has no User-Agent header. | Set a User-Agent header on the request, or contact support if it persists. |
PagerDuty answers 200 but no alert appears. | The PagerDuty webhook subscription does not include the incident events. | In PagerDuty, add the incident events to the subscription, including incident.triggered. |
| CloudWatch or CloudTrail alerts never appear, and the SNS subscription shows “pending confirmation”. | SRE Agent could not confirm the subscription, and SNS does not retry on its own. | Request the confirmation again from AWS. A 502 in the delivery log means the confirmation call failed. |
| A resolved notification creates nothing. | A resolve for an alert SRE Agent never had open is accepted and ignored. This is normal, for example for a CloudWatch alarm’s first INSUFFICIENT_DATA to OK change. | No action needed. |
Posting the same alert to the API again answers 200, not 201. | Alerts are matched by fingerprint, so a repeat updates the existing alert. | To record separate alerts, send your own fingerprint or vary source_id. |
| Alerts keep arriving past your monthly allowance. | Alerts are never refused at the limit. They are counted. | Check usage on the Subscription page. |
An investigation does not start
Open the alert. When an alert is skipped, its timeline says “Not investigated” followed by the reason.| Reason shown | Cause | Fix |
|---|---|---|
| The alert is muted. | Someone silenced it. | Unmute the alert, or start an investigation by hand. |
| Somebody acknowledged the alert. | A person took ownership. Automatic investigation resumes if it resolves and fires again. | Start one by hand for a second opinion. |
| Below the automatic-investigation severity floor. | The alert’s severity is too low for automatic investigation. | Start one by hand, or raise the alert’s severity at the source. |
| The alert is flapping. | It repeatedly fires and clears on its own, and another investigation would find the same thing. | Tune the rule that produces it, or start one by hand. |
| The same incident is still open, or the alert was investigated recently. | Nothing resolved since the last investigation, so there is no new occurrence to diagnose. | Open the linked earlier investigation. Start a new one by hand if you want one. |
| The alert reached its daily or weekly investigation cap. | It fires so often that automatic investigation is capped. | Tune the rule, or start one by hand. |
| The organization is over its investigation budget for the moment. | A burst of alerts used the short-term budget. | Wait for it to recover, and the alert is investigated again. |
| The alert says it was not investigated because the organization has spent its monthly AI budget, so automatic investigations are paused. | The organization reached its monthly AI budget. Alerts are still received and notified. | Raise the budget on Usage & Spend with Change budget (Set a budget when none is set), or wait for the month to reset. You can still start investigations by hand. |
| The Investigations page shows an upgrade banner. | You used your plan’s monthly investigation allowance. | Wait for the next month or change plans. See Limits. |
A notification or page was not delivered
The delivery record for the page names why it was skipped or failed.| You see | Cause | Fix |
|---|---|---|
| ”no push device” | The person has not enabled push on any device. | On Account, enable push on the device. |
| No Enable button on an iPhone. | The app is not on the Home Screen, or iOS is older than 16.4. | Add the app to the Home Screen and update iOS. |
| ”no Pushover user key” | The person has not saved a Pushover user key. | Save the key on Account. |
| A failure naming the Pushover user | Pushover refused the saved key. | Save the key again on Account. |
| ”no Telegram chat linked” | Telegram is not linked, or /stop was sent to the bot. | Press Link Telegram on Account, Telegram, and press Start in the bot. |
| ”not linked to Slack” | The person has not linked their Slack account. | Run /sre-link <your email> in Slack. |
| ”this organization has no Slack workspace” | No Slack workspace is connected. | Connect Slack on Integrations. |
| ”email notifications are turned off” | The person opted out of email for non-page notifications. | Turn email notifications back on, or add another channel. Pages are never blocked by this opt-out. |
| A PagerDuty target says the alert came from PagerDuty. | An alert that PagerDuty reported is never paged back to PagerDuty. | Page it through a Slack, webhook or on-call target, or handle it in PagerDuty. |
A runbook step is refused
| You see | Cause | Fix |
|---|---|---|
| ”No enabled connector found for type” | The organization has no enabled connector of that type. Data sources do not count, connectors are separate. | Add and enable a connector of that type. |
| ”The runbook is not approved” | Only approved runbooks run. A draft is refused. | Approve the runbook. A dry run works on a draft. |
| ”Unresolved reference(s) in action” | A {{...}} reference points at a step or field that does not exist. | Fix the reference, or add the step it names. |
| A restart step needs a resolved deployment name. | The step’s resource is empty, a reference that did not resolve, or a * that no discovery step filled in. | Fill in the name, or add a discovery step before it. |
| The target is unavailable. | The target is reached through a remote agent that is offline. The step refuses rather than running outside the agent. | Bring the remote agent back online. |
AccessDenied from AWS | The connector’s role lacks the action the step needs. | Add the action to the role. |
| A write step is refused above 50 items. | Discovery found more items than a write step may act on at once. | Narrow the discovery step. See Limits. |
| A step waits for approval. | The runbook’s Who approves setting asks a person, or an approval gate step is reached. | Approve the run, or change Who approves if no one should be asked. |
A Slack message is missing
| Symptom | Cause | Fix |
|---|---|---|
| The bot does not post in a private channel. | A bot cannot join a private channel by itself. | Run /invite @SRE Agent in the channel. Public channels need nothing, the app joins on the first post. |
| No direct messages, private-channel replies or reactions work. | Your workspace installed the app before those permissions existed. | Choose Add to Slack again to re-authorize. |
| The bot says to link your account. | Questions and commands need a Slack account linked to an SRE Agent account in your organization. | Run /sre-link <your email>. |
| Replies in a thread do not become card comments. | Only replies in the thread of an alert the app posted are captured, and only on a card that names that alert. A reply that mentions the bot is a question and is answered in the thread. | Link the thread to the card, or add the comment on the card. |
| A card moved but nothing posted in Slack. | A card posts to Slack only when it is linked to a thread. | Link a thread when you open the card, or use the alert’s thread. |
| A fix outcome did not post where you expected. | /sre-fix replies in the channel you ran it from. Requests without a channel post to the Remediation Channel under Settings, Notifications (in the Slack section), or to the Info channel when that is blank. | Set the Remediation Channel. |
| A reaction does nothing. | Only the :eyes: and :rotating_light: reactions on alert messages the bot posted act, and only for a linked account. | React on the alert message from a linked account. |
| The Create a board card step is missing in Workflow Builder. | The step needs your own SRE Agent Slack app on a paid Slack plan. | Use the Send a web request step with the ops board webhook. |
A fix request is declined or refused
A fix request that is accepted can still end as a decline. Open Control Tower to see each request move from analyzing to open, declined or failed, with its reasoning attached.| Symptom | Cause | Fix |
|---|---|---|
The webhook answers 422 with a reason. | The request was refused before analysis. The reason names the problem. | Fix what the reason says, then resend. |
| The reason says the GitHub App is not installed. | Fix pull requests need the GitHub App. | Install the App from Integrations, GitHub. |
| The reason says your plan does not include fixes. | Fix pull requests need the Business plan. | See Plan matrix. |
| The target matches no mapped service or repository. | The service or repo you named is not one the App can reach or you mapped. | Use a listed alias or owner/name. The reason lists what is mapped. |
| A short repository name is ambiguous. | Two mapped repositories share the name. | Use owner/name. |
| The repository already has 5 open bot pull requests. | The per-repository limit is reached. | Merge or close an open one. |
| A new fix request is refused with “5 requests are already being analyzed; name a service or repo, or wait for one to finish”. | 5 requests that name no service or repo are already being analyzed, and this one names none either. | Name a service or repo in the request, or wait for one of the 5 to finish. |
| A named file is refused. | The file is outside the area the agent may change, or under .github/. | Name files inside the allowed area, or leave files out. |
| The request shows “waiting for a person”. | The repository is set to Wait for a person. | Press Send on the card, or in the request’s drawer in Control Tower. |
| A Jira or Zoho Sprints issue is ignored. | It has no sre-agent-fix label or tag, or was refused for one of the reasons above. The reason is in the response and, when commenting is set up, on the issue. | Add the label or tag and fix the reason. |
| The request was declined with reasoning. | The agent found no safe change, or no mapped repository fit. | Read the reasoning in Control Tower. Add detail and send again with the same dedup_key to replace it. |
| An earlier pull request closed on its own. | Sending the same ticket or dedup_key again replaces the earlier request and closes its pull request. | Use a new key when you mean a separate request. |
An SLO shows degraded or no data
| Symptom | Cause | Fix |
|---|---|---|
| The SLO is Degraded and shows a reason. | Its data source is unreachable, or the source answered with nothing at all for the window. SRE Agent reports no compliance number rather than a false 100%. | Fix the data source. The SLO recovers on its own once a window has data, within a few minutes. |
| An idle service shows 100%. | The source counted zero events, so nothing violated the objective. | No action needed. |
| A Loki SLI query is refused. | The query is a log stream selector, not a metric query. | Wrap the selector in count_over_time or rate and sum it. |
| A CloudWatch Logs query is refused. | The query does not end in a stats aggregation. | End the query with stats. |
| The SLO reports that its CloudWatch Logs data source “has no log group to search, so no query was sent”. | The CloudWatch Logs data source has no log group to search. | Fill in Default log group under Settings, Data Sources. |
| The SLO reports that its New Relic data source “has no account to query, so no query was sent”. | The New Relic data source has no account to query. | Fill in Account ID under Settings, Data Sources. |
| A synthetic SLI waits for its first run. | The check has not run yet. | Wait for the first scheduled probe. |
| The query type is not supported. | The SLI uses a query type outside the supported list. | Edit the SLI to use a supported type, or deactivate it. |
| The SLO wizard redirects to the SLO list. | The wizard is not on your plan. | See Plan matrix. |
The deploy gate is locked
GET /api/deploy-gate/<service> answers 200 when a deploy is safe and 423 Locked when it is not. The response lists each cause under reasons, and lists things you should know about under advisories.
A deploy is also blocked when SRE Agent cannot measure a service’s reliability. An SLO whose data source is unreachable is Degraded, and a degraded SLO holds its service’s gate.
| Way out | How |
|---|---|
| Fix the data source. | The gate clears on its own within about five minutes. This is the intended path. |
| Suspend the SLO. | On the SLO dashboard, suspend it. A suspended SLO never blocks, and is reported under advisories instead. Members and above can do it, and Resume reverses it. |
| Stop the SLO from gating. | Set the SLO’s Gate deploys on this SLO control to Never gate deploys. |
| Override for an emergency. | POST /api/deploy-gate/<service>/override allows deploys for 1 to 72 hours, 4 by default. The override is attributed to your API token and recorded in the audit trail. |
You reached a plan limit
| Symptom | Cause | Fix |
|---|---|---|
| ”You’ve reached your plan’s limit of N” | You hit a count limit such as users or data sources. | Remove an item or change plans. |
| ”This feature isn’t included in your current plan.” | The feature needs a higher plan. | See Plan matrix. |
Related
- Triage alerts: how alerts are handled.
- On-call notifications: how pages reach people.
- Work incidents from Slack: Slack setup and commands.
- Gate deploys on reliability: gate overrides and thresholds.

