<!-- Canonical: https://www.agentlist.io/learn/guides/evaluate-an-agent -->

[The agent fieldbook](https://www.agentlist.io/learn/guides)

# Find out whether an agent actually helps

Build a small, honest evaluation from your own work. Score observable outcomes, include your review time, and diagnose failures before changing the prompt or buying a different tool.

Chapter 09 of 1015 min readIntermediate

[Jump to the exercise](#try-it)

Measure the result, one useful step at a time.

Keep this chapter open alongside your next real task.

## What you’ll be able to do

-   Build a representative test set with an explicit pass condition.
-   Compare agents using complete outcomes and total human effort.
-   Turn a failed run into a targeted improvement and regression check.

## 1\. Evaluate a job you recognize

Choose a task you understand well enough to judge. Examples include extracting fields from invoices, drafting release notes from merged changes, or finding evidence for a short research brief. A spectacular demonstration on someone else's task does not tell you how much supervision your work will require. Your first evaluation can be a spreadsheet with a dozen representative examples.

Include ordinary cases, difficult cases, and cases where the right response is to stop or report missing information. Preserve the same starting materials for each candidate. Remove private data that is unnecessary for the test. Write down the product, model if visible, tool access, instructions, and date. You are evaluating that complete setup. A model name alone does not capture the connections, permissions, or workflow that produced the result.

Worked example

### A small test set

Illustrative invoice task: six clear invoices, three with unusual layouts, two with a missing purchase-order number, and one credit note. The output must preserve document type, supplier, date, currency, total, and source filename. Missing values must remain explicitly unknown.

## 2\. Define passing before viewing the answers

Write criteria that another person could apply. ‘Looks good’ rewards confidence and formatting. ‘Every total matches its source, every currency is explicit, and missing values are not invented’ tests the actual job. Separate essential conditions from quality preferences. A well-written brief with an unsupported central claim should fail the evidence requirement even if its prose earns a high style score.

For actions, inspect the resulting state. Anthropic's evaluation guidance distinguishes an agent's transcript from the final outcome in its environment. If the task is to create a draft, verify that the draft exists in the correct location. If the task is to edit a record, compare the saved record with the requested values. The agent's final statement is useful as a report, but it is not an independent check of its own work.

1.  ### Set the essential conditions
    
    Name the errors that make a result unusable, such as a wrong currency, missing source, or unauthorized send.
    
2.  ### Define partial credit
    
    Use a small rubric for useful but incomplete results. Explain exactly what each score means.
    
3.  ### Choose the evidence
    
    Specify the source comparison, saved artifact, or destination state that proves each condition.
    

### Keep one row per attempt

Record the task, candidate, model, starting inputs, accepted or rejected result, reason, human minutes, elapsed minutes, usage cost, and evidence links. Include retries and failed attempts. Keep one-time setup separate so a reusable environment does not disappear from the cost calculation.

Try it yourself

## Calculate the value of a completed task

Adjust completion rates, review time, and task costs in an illustrative evaluation. See how the conclusion changes when rework and failed attempts count.

Tasks attemptedTasks acceptedManual minutes per accepted taskHuman minutes per attemptAgent cost per attempt ($)Human time value ($/hour)

#### Your sample, measured

**Acceptance rate**: 70%

**Net human time saved**: 190 min

**Cost per accepted result**: $14.29

**Estimated net value**: $150.00

Time saved = accepted tasks × manual minutes − all attempts × human minutes. Net value converts that time to dollars and subtracts agent costs. Include review, prompting, and recovery in human minutes.

Illustrative planning model, not a benchmark. Unfinished work earns no time saving. This excludes setup, infrastructure, and downstream errors; assess those separately.

## 3\. Keep the comparison fair

Run candidates on the same examples with the same authorized access and comparable instructions. Keep a clean starting state so one candidate does not inherit another's completed work. Save the output and enough activity history to explain failures. Record any human intervention, including clarifying a request, correcting a tool target, or manually finishing a step.

Repeat a few important examples to see whether success is consistent. A single pass shows that the setup can succeed once. It does not establish a dependable success rate. If you adjust the prompt after every failed example, you are developing the workflow, not measuring its performance on unseen work. Reserve several examples for a final check after the instructions settle. For subjective writing comparisons, hide the candidate names and review outputs side by side against the same rubric.

### Keep the sample claim modest

Ten successful tasks out of twelve is an observation about twelve tasks. It is useful pilot evidence, not proof that future work will succeed at the same rate.

## 4\. Count review, recovery, and failed attempts

Record how long the same work takes without the agent, then measure your preparation, review, and correction time with the agent. Also record elapsed waiting time if it affects delivery. An agent that finishes in two minutes but needs fifteen minutes of checking may still be useful, but the time-saving claim must include those fifteen minutes.

Use cost per accepted result when comparing paid runs. Add the direct usage cost of successful and failed attempts, then divide by accepted outputs. Track human time separately or convert it using an explicit hourly value. The calculator uses illustrative inputs so you can see which assumptions control the result. It cannot establish actual vendor prices or your future failure rate. For a pilot, keep the raw observations next to every calculated figure.

Worked example

### A worked comparison

Suppose a manual task takes 30 minutes. An agent-assisted version needs 5 minutes of setup, 8 of review, and 7 of correction. Your hands-on saving is 10 minutes. If a separate failed attempt needs another 12 minutes of recovery, include that attempt when calculating the average across the batch.

## 5\. Fix the cause of the failure

Read failed runs from the point where reality and the task first diverged. Did the agent lack a source, misread a field, choose the wrong tool, lose a required constraint, or claim completion without checking? Those causes suggest different changes. Adding stronger wording to the prompt will not repair an expired connection or supply a missing document.

Change one meaningful variable, then rerun the failed example and a few previously successful examples. If the fix requires a new instruction, make that instruction observable: ‘Leave a missing purchase-order number blank and add missing\_purchase\_order to the issues field.’ Keep the failure as a future test case. Google documents comparing model-generated judgments with human ratings; apply the same principle if you use an agent to grade outputs. Check whether its verdicts agree with your criteria before relying on its scores.

Worked example

### From failure to regression check

Failure: a credit note was counted as a new invoice. Root cause: the output schema had no document-type field. Change: add invoice or credit note, with a source check. Retest the credit note and ordinary invoices to ensure the fix preserves correct totals.

## 6\. Choose the scope that the evidence supports

An evaluation can justify a narrower deployment than you originally imagined. The agent might be reliable for extracting ordinary invoices while unusual layouts still need manual handling. Define a routing rule for those exceptions and test whether the agent recognizes them. Do not hide difficult cases by removing them from the denominator after the run.

Write a short decision record: task, test set, pass criteria, results, human effort, known failures, and allowed next use. Keep the baseline examples so you can rerun them after a model, tool, prompt, or workflow change. Review new real-world failures and add representative cases. The purpose of the evaluation is a practical decision about where the agent earns its place and where the current setup still needs a person.

## When things go wrong

Start with the failure you can observe.

Every agent gets a high score, but none is useful.

Replace style-heavy criteria with essential task outcomes. Include the time needed to verify and repair the result.

Results improve only on examples you keep rerunning.

Separate development examples from a held-back check set. Freeze the instructions before running the check set.

An agent grader approves obvious mistakes.

Compare its ratings with human judgments on known passes and failures. Revise the rubric and retain manual checks for criteria the grader cannot assess reliably.

## Put it into practice

0 / 5

Use this checklist on your next real task.

The examples represent ordinary work, exceptions, and missing information.Essential pass conditions were written before inspecting candidate outputs.Completion is checked against artifacts or destination state.The results include human intervention, rework, and failed attempts.The decision names a supported scope and preserves regression examples.

Your checklist is saved in this browser. Exercise inputs are not saved.

## Sources & further reading

Primary references for the ideas in this chapter. Product behavior can change; check the documentation for the version you use.

-   [Anthropic: Demystifying evaluations for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)
    
    Supports evaluating environmental outcomes, preserving execution evidence, and interpreting repeated trials rather than relying on final narration.
    
-   [Google Cloud: Evaluate a judge model](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/evaluate-judge-model)
    
    Explains comparing automated evaluation judgments against human ratings. This guide's pilot rubric and arithmetic are illustrative.
    
-   [Google Cloud: View and interpret evaluation results](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/eval-python-sdk/view-evaluation)
    
    Documents per-example and aggregate evaluation results and the distinction between scoring one output and comparing a pair.
    

Edited October 7, 2026 · Examples are illustrative unless attributed.

---

Source: [Find out whether an agent actually helps | agentlist.io](https://www.agentlist.io/learn/guides/evaluate-an-agent). This is the public page rendered as Markdown; interactive controls require the website.
