Taxonomy — 28 capabilities · 112 systems

Agent capabilities

Capabilities are atomic tags that describe what an agent can actually do — run code, edit files, drive a browser, take phone calls — independent of which product category it belongs to. A CLI agent and a cloud agent can share 'long-horizon tasks'; a voice agent and a digital employee can share 'human handoff'. The tags let you compare across categories rather than only inside them.

Tags verified against the catalog seed of 2026-09-25. Filter the agent index by capability to compare across categories.

F.01

Agent capabilities — What the agent itself does

Background runs /background-tasks

Asynchronous cloud or background execution you can walk away from.

10 agents →

Browser operation /browser-ops

Drives a real web browser: navigate, click, fill, extract.

12 agents →

Cited answers /citations

Returns answers grounded in identifiable sources.

08 agents →

Code execution /code-exec

Runs code in a real environment or sandbox.

16 agents →

Code review /code-review

Reviews diffs and pull requests with feedback.

04 agents →

Desktop control /desktop-use

Pixel-level computer use — screenshots, mouse, keyboard.

03 agents →

File editing /file-edits

Reads and writes project files, including multi-file edits.

35 agents →

Git & PRs /git-pr

Branches, commits, and pull-request workflows.

09 agents →

Human handoff /human-handoff

Escalates, transfers, or requests approval from a human.

28 agents →

Incident response /incident-response

Alert, on-call, and incident-triage loops.

07 agents →

Long-horizon tasks /long-horizon

Unattended multi-step work running minutes to hours.

26 agents →

Model routing /model-routing

Selects or routes across multiple underlying LLMs.

13 agents →

Persistent memory /memory

Remembers context across sessions or runs.

12 agents →

Repo context /repo-context

Indexes or reasons over a whole repository, not just open files.

34 agents →

Shell access /shell

Executes terminal/shell commands.

25 agents →

Voice I/O /voice

Realtime voice or telephony input/output.

09 agents →

Web research /web-research

Searches, browses, and synthesizes information from the web.

17 agents →

Workflow builder /workflow-builder

Visual or no-code composition of agent workflows.

06 agents →
F.02

Infrastructure capabilities — How it composes into a stack

FAQ

Frequently asked

What are AI agent capabilities?

Capabilities are atomic tags that describe what an agent can actually do — run code, edit files, drive a browser, take phone calls — independent of which product category it belongs to. A CLI agent and a cloud agent can share 'long-horizon tasks'; a voice agent and a digital employee can share 'human handoff'. The tags let you compare across categories rather than only inside them.

How is this different from the category taxonomy?

Categories answer 'what kind of product is it' — an AI IDE, a cloud agent, a voice agent. Capabilities answer 'what can it do' — execute code, call MCP tools, route between models. Categories partition the catalog; capabilities cut across it.

What do the two facets mean?

Agent capabilities are things the agent itself does — file editing, web research, voice I/O. Infrastructure capabilities describe how the system composes into a stack — API-native, self-hostable, agent payments, sandboxed runtimes. Most infrastructure entries are building blocks rather than end-user agents.

How are capability tags assigned?

Each system is tagged only with capabilities verifiable from its documented feature set in the catalog seed. Tags are not self-reported by vendors and are re-checked when the catalog is regenerated.