Skip to main content
SRE Agent is a hosted incident platform. Your monitoring tools send it alerts, an AI agent investigates each one against your own metrics and logs, and the result lands where your team already works: a Slack thread, the dashboard, a board card, and when you want one, a pull request that fixes the problem. You do not install anything. You point your alert source at a webhook URL, connect the data the agent should read, and the first investigation starts when the next alert fires. The quickstart takes you there in one sitting.

The incident path

1

An alert arrives

Grafana, PagerDuty, Datadog, New Relic and CloudWatch (through SNS) each get their own webhook URL. For anything else you can post to the REST API with an API key. SRE Agent normalizes every alert to one shape with a severity, labels and a description, and posts it to Slack if you connected Slack.
2

The agent investigates

For each new alert the agent queries the data sources you connected, such as Prometheus, Loki, CloudWatch, Elasticsearch, Datadog and New Relic. It works through the metric that fired, resource usage and recent changes, correlations with past incidents, and then forms and checks hypotheses about the root cause.
3

You get a root cause and next steps

Open the investigation to read the summary, the root cause, the recommendations, the evidence behind each finding and a timeline of every query the agent ran. You can ask follow-up questions from Slack.
4

Work lands on the board

You can file a card on the board straight from an investigation. On the Business plan, you can also file a ticket in Jira, Zoho Sprints or GitHub.
5

A fix arrives as a pull request

On the Business plan, with the GitHub App installed, SRE Agent can open a pull request for the fix. The pull request is the approval step: nothing changes in your repository until your team merges it.

What else is in the product

Alongside the investigation flow, SRE Agent covers the rest of running an on-call rotation:
  • On-call: daily, weekly or custom rotations, overrides, and escalation. Pages reach people by push notification from the installed web app, Pushover (which can break through Do Not Disturb), Telegram, Slack direct message and email. There is no SMS or phone call. You can also forward pages to PagerDuty or a webhook.
  • SLOs and synthetic checks: define service level objectives on your own metrics and probe endpoints from outside.
  • Runbooks and automations: multi-step procedures with approval gates, started by hand, by an alert pattern or on a schedule.
  • Status page: a public page for your customers with components and incidents.
  • Control Tower: an agent that watches your estate on a schedule, cuts alert noise and writes up problems no alert covers as findings you can act on. It is on the Business plan.
  • Security and FinOps: findings and cost analysis for the infrastructure you connect.

Who it is for

On-call engineers who want the first ten minutes of an incident done before they open a laptop, and team leads who want every incident to leave a written root cause behind.

Plans

Every account starts on the Free plan, with no payment step. Pro and Business start with a 14-day free trial that you begin by upgrading from inside the app. You can cancel any time from Manage billing. Support and contract terms are on the pricing page. Each guide in this documentation shows a plan badge when a feature needs more than the Free plan.
Investigations run on SRE Agent’s shared AI until you add your own provider. Adding your own is optional on every plan and is the way to keep your alert text, logs and hostnames inside credentials you control. The quickstart shows where.