Choose your first AI agent
Start with a job you can judge. Find the right kind of agent, run a small trial, and choose with evidence you can actually use.
Jump to the exerciseWhat you’ll be able to do
- Turn a vague interest in agents into a testable job.
- Compare candidates against your own work and constraints.
- Recognize when a simpler tool is the better choice.
1. Start with a job you already understand
Pick something you have done before: compare three products, repair a small bug, prepare a meeting brief, or organize a folder of reference documents. Familiar work gives you a useful advantage. You know what a good result looks like, where mistakes tend to hide, and how much effort the task normally takes. Your first agent should earn trust on that familiar ground.
Write the job as an observable result. ‘Help with research’ leaves the finish line open. ‘Produce a comparison of three scheduling tools using their current official documentation, with an evidence link for each requirement’ creates a result you can inspect. Keep the first task small enough to review in one sitting. A successful trial should teach you something about the agent without creating a second project to supervise.
Illustrative first job
You normally spend 40 minutes preparing for a customer interview. Give the agent a public company website and a short interview objective. Request a one-page brief with five sourced facts, five interview questions, and a separate list of unknowns. Review the facts before using the brief.
2. Match the agent to where the work happens
An agent needs a way to reach the materials and tools involved in your task. A research agent needs suitable search and source access. A coding agent needs the project, its instructions, and a way to run relevant checks. A browser agent needs access to the correct website and account. A workflow agent needs the particular applications involved in the handoff. A broad product description does not establish that these connections are available to you.
Draw a short route from input to output. For an interview brief, that might be public pages, a synthesis step, and a document. For a support workflow, it could include a ticket, an account record, a draft response, and a human review. Circle every place where access is required. Check those connections before comparing writing quality or clever demonstrations.
For coding work, choose the working surface as well as the product. An editor keeps frequent code and visual feedback close. A terminal suits a task whose acceptance checks run as commands. A remote run needs a reproducible environment and a reviewable handoff. One product may offer several surfaces; test the one you intend to use.
Find the agent for your next task
Choose your task, working environment, and review style. Get a starting category and a first trial to run.
Start with an editor-based coding agent
Your job needs repository access, file edits, and a way to run the relevant checks.
Keep the first task short enough to watch and interrupt.
Your first trial: Fix one reproducible bug in a disposable project copy. Compare the diff and test output with the original failure.
Browse this category This is a starting category, not a product ranking. Verify each candidate against your task.3. Decide how you will judge the result
Set a few criteria before running a trial. Include correctness, completeness, evidence quality, and the effort needed to repair the result. Add a task-specific criterion that would matter in real use. A coding change might need to preserve keyboard navigation. A research brief might need to distinguish a vendor promise from a capability that you tested. A document workflow might need to keep the original files intact.
Separate requirements from preferences. If the agent cannot access an essential source under your organization's rules, attractive prose does not compensate. Treat that as a failed requirement. Among candidates that meet the requirements, compare conveniences such as export format, response speed, and how clearly they explain an incomplete run. Give each result a short written reason rather than a mysterious overall score.
Set the pass conditions
Write three observable requirements, such as five accurate facts, working source links, and no invented company details.
Set a review budget
Decide how long you can spend checking and repairing the output before the trial loses its practical value.
Set a stop condition
Stop the trial if the agent requests unnecessary access, repeats the same failure, or exceeds your chosen time or usage limit.
4. Give two candidates the same small trial
Use the same inputs, instructions, and acceptance criteria for each candidate. Save the initial brief so that an accidental change does not make the comparison unfair. Record the setup you used, including relevant permissions and whether browsing or other tools were enabled. If one candidate needs a different workflow, include the setup time in your notes.
Run a normal case and one awkward case. For the interview brief, the awkward case could be a company with sparse public information. Watch whether the agent clearly reports the gaps or fills them with plausible detail. You are testing how the product behaves when the world is inconvenient. Do not infer reliability from one polished answer. Repeat a representative task before making the agent part of a recurring process.
5. Count the work around the answer
Measure the full task: preparing inputs, explaining context, waiting, checking, correcting, and moving the result into the place where you need it. An answer that appears in two minutes can still require half an hour of cleanup. Conversely, a slower run may be useful if it produces an artifact you can inspect and use with little repair. Keep cost and elapsed time separate from your own attention.
Read the current plan and usage terms directly before subscribing. Verify the features, limits, and account requirements that matter to your trial. Record the date because those details can change. If your work includes confidential material, check the applicable data handling and administrative controls before uploading it. Public sample data can establish basic fit while you resolve those requirements.
Illustrative comparison
Candidate A finishes in four minutes but takes 18 minutes to fact-check and repair. Candidate B finishes in seven minutes and takes six minutes to review. If both meet your requirements, B uses less of your attention for this job. That result does not prove B is better for every task.
6. Choose a bounded next step
Choose the candidate whose results and working conditions fit the job. Save the winning brief, one accepted result, and a short note about the failures you observed. For the next week, use the agent on the same family of tasks. Expand to adjacent work only when you can describe what new access, judgment, or verification that work requires.
You can also decide that the task does not need an agent. A template may solve a repeated writing problem. A saved search may answer a narrow research question. A conventional automation may suit an exact sequence. The useful outcome of the trial is a better way to do the work. Keep your evidence small and concrete so that changing products later does not require rebuilding your entire process.
When things go wrong
Start with the failure you can observe.
Every agent looks equally impressive.
Replace generic demos with one task from your own work. Include a detail that requires checking a source or preserving an existing constraint.
The trial takes longer than doing it myself.
Identify the expensive stage. If context setup dominates, reuse a short brief. If corrections dominate, reduce the task or try another candidate.
The agent succeeds once and fails next time.
Repeat with saved inputs and settings. Record which requirement varies. Keep human review on that requirement until you have a more dependable process.
Put it into practice
0 / 5Use this checklist on your next real task.
Your checklist is saved in this browser. Exercise inputs are not saved.
Sources & further reading
Primary references for the ideas in this chapter. Product behavior can change; check the documentation for the version you use.
- Anthropic: Building effective agents
Supports the workflow versus agent distinction and the tradeoff between added autonomy, cost, and complexity. The trial method here is an original practical exercise.
- NIST: AI RMF Playbook
Provides the broader govern, map, measure, and manage approach behind evaluating an AI system in its actual context of use.
Edited October 7, 2026 · Examples are illustrative unless attributed.