ARTICLE 07 · EVALUATION · agentlist.io · 2026-10-07 · 6 min read
A coding-agent trial you can actually score
Give two candidates the same small bug. Count the work you do after the answer arrives.
Pick a bug whose failure you can reproduce and whose correction you can recognize. A good first trial is small enough to inspect in one sitting. Avoid a broad request such as “clean up the codebase,” where a large diff can obscure whether anything useful happened.
Freeze the starting conditions
Save the starting commit, brief, relevant tool versions, model choice, and permission settings. Use a clean worktree for each attempt. Keep existing uncommitted work out of the trial. Give both candidates the same inputs, then record any extra hints or intervention you provide.
Run the reproduction yourself first. Confirm that the expected behavior is correct. If a required service is unavailable to both candidates, supply a suitable test environment before comparing results.
Copy this brief and fill in the blanks
Task: Fix [observable failure] in [component].
Starting point: [commit] in a clean worktree.
Reproduce: [steps or command, input, actual result].
Expected result: [observable correct behavior].
Trace the cause before editing. Preserve unrelated work.
Keep the change within [scope].
Run [relevant check] and check [important edge case].
Stop at [time or usage limit], or if [specific blocker].
Do not deploy or change external data for this trial.
Return the patch, cause, checks run and their results,
and anything that remains unverified.Use pass conditions before preference scores
Accept the result only if it fixes the stated behavior, stays within scope, and passes the relevant checks. A failed requirement is still a failure when the explanation is clear or the answer arrives quickly. Record the reason for rejection so the next trial tests a useful question.
- Correctness
- Does the original reproduction now behave as expected? Does the awkward case still work? Save the evidence.
- Scope
- Does every changed file support the task? Note unrelated edits or changes that hide the failure.
- Verification
- Which checks actually ran on the final patch? Distinguish passing, failing, and not run.
- Human effort
- Record minutes spent preparing, prompting, reviewing, repairing, and recovering. Keep waiting time separate.
- Cost
- Record usage cost and setup expense. Include failed attempts in cost per accepted result.
Keep a record for every attempt
A spreadsheet row is enough: task, candidate, model, starting commit, accepted or rejected, reason, human minutes, elapsed minutes, usage cost, and evidence links. Record retries as attempts. Do not retain only the run you liked.
For an illustrative calculation, suppose five attempts produce three accepted results. You spend 50 human minutes across all attempts. Doing those three tasks manually would have taken an estimated 90 minutes. The observed saving is 40 human minutes, before one-time setup. If total usage cost was $6, usage cost per accepted result is $2. These sample numbers describe no particular product.
Track the two unfinished tasks separately: they still need work. With no accepted results, report zero accepted results and total cost; cost per accepted result is undefined.
Decide what to try next
Repeat with another representative task before a recurring commitment. If both candidates fail because the brief is ambiguous, improve the brief. If the environment is broken, repair it. If one repeatedly meets your conditions with less review work, try it on the next bounded task.
The fieldbook’s evaluation chapter covers acceptance rate, cost, and recovery in more depth. Its downloadable worksheet gives you a place to record a trial.