12 · ops-agent · 07 systems

On-call & ops agentsAgents that hold the pager — triage, investigate, resolve.

Ops agents sit inside the incident and security loop: they watch alerts, correlate telemetry, investigate pages, draft root-cause analyses, and propose or apply remediations. AI SREs handle production incidents; AI SOC analysts (Dropzone, Simbian) handle security queues.

SEC.01

Systems in this category

Cleric

Cleric · ops-agent

Cleric is a dedicated AI SRE: it investigates production alerts across logs, metrics, deploys, and code history, then hands on-call engineers a root-cause hypothesis with evidence.

chat/api · cloud · Enterprise (quote-based)

Resolve

Resolve · ops-agent

Resolve (resolve.ai) is an AI SRE from ex-Splunk/Observability leadership: it triages alerts, investigates incidents across your stack, and drafts remediations for approval.

chat/api · cloud · Enterprise (quote-based)

incident.io

incident.io · ops-agent

incident.io is the incident-management platform whose AI assistant drafts timelines, summarizes channels, surfaces similar past incidents, and nudges responders — inside the tool that already runs your incidents.

chat/api · cloud · Free tier; per-seat platform pricing (Response/Assist tiers)

Rootly

Rootly · ops-agent

Rootly is the incident-response platform with AI woven through it: auto-drafted summaries, suggested next steps, retrospectives, and noise reduction on the alert stream.

chat/api · cloud · Per-seat plans; enterprise tiers — verify

Bits AI

Datadog · ops-agent

Bits AI is Datadog's in-platform agent: it answers observability questions in natural language, investigates anomalies, and drafts incident summaries against the telemetry you already pay Datadog for.

chat/api · cloud · Included in Datadog plans; usage-based add-ons — verify

Dropzone AI

Dropzone AI · ops-agent

Dropzone AI is the AI SOC analyst: it investigates every alert end-to-end — triage, evidence gathering, verdict — so human analysts only see the ones that matter.

chat/api · cloud · Enterprise (quote-based)

Simbian

Simbian · ops-agent

Simbian builds autonomous SOC agents that hunt, investigate, and respond across the security stack, plus GRC automation — aimed at teams running lean against enterprise-scale alert volume.

chat/api · cloud · Enterprise (quote-based)
SEC.02

Who it's for

SRE, platform, and security teams buried in alerts — the 3am triage, the first-hour investigation, the ticket hygiene — who want an agent to take the first pass before a human decides.

SEC.03

Tradeoffs

The category is trust-constrained: agents get read (and sometimes write) access to observability and cloud accounts, most deployments still gate remediation behind human approval, and pricing is enterprise-contract territory.

FAQ

Frequently asked

What is an ops agent?

Ops agents sit inside the incident and security loop: they watch alerts, correlate telemetry, investigate pages, draft root-cause analyses, and propose or apply remediations. AI SREs handle production incidents; AI SOC analysts (Dropzone, Simbian) handle security queues.

Who should use an ops agent?

SRE, platform, and security teams buried in alerts — the 3am triage, the first-hour investigation, the ticket hygiene — who want an agent to take the first pass before a human decides.

What are the tradeoffs of ops agents?

The category is trust-constrained: agents get read (and sometimes write) access to observability and cloud accounts, most deployments still gate remediation behind human approval, and pricing is enterprise-contract territory.

Which ops agents are in the catalog?

7 as of 2026-09-25: Cleric (Cleric), Resolve (Resolve), incident.io (incident.io), Rootly (Rootly), Bits AI (Datadog), Dropzone AI (Dropzone AI), Simbian (Simbian).

What is an AI SRE agent?

An agent that does the front half of incident response: it watches alerts, pulls logs/traces/metrics, forms a hypothesis, and drafts a root-cause summary or remediation for an on-call engineer to approve. Cleric and Resolve are the dedicated entrants; Datadog Bits AI ships inside an existing observability platform.

How is an ops agent different from a cloud coding agent?

Cloud coding agents take a task and return a pull request. Ops agents take an alert and return an investigation — timelines, correlated signals, suspected cause — and increasingly a suggested fix. The output is a decision, not a diff.

Do ops agents remediate autonomously?

Mostly no, by design: current deployments emphasize investigation and recommendations with human approval gates before actions like rollbacks or scale-ups. Autonomous remediation exists but is typically scoped to low-risk playbooks.