The agent fieldbook

Turn one good result into a workflow you can trust

Capture the parts that worked, define the handoffs, and test recovery before adding a schedule. Build a small operating system for one recurring job, with an owner and a clear finish.

Jump to the exercise
The little railway that gets things donereview, then repeata clear handoffarrive somewhere
Make it repeatable, one useful step at a time.

What you’ll be able to do

  • Describe a recurring task as explicit inputs, stages, and acceptance checks.
  • Design a restart path that preserves completed work and avoids duplicates.
  • Decide which parts should run automatically and which need review.

1. Start from a task that already worked

Choose a recurring job that you have completed successfully with an agent at least once. Save the brief, source set, useful intermediate artifacts, final output, and the corrections you made. Those corrections reveal work that the first prompt did not capture. A workflow copied only from the polished final answer misses the decisions that made the answer usable.

Describe the job in one sentence with a bounded outcome. For example: ‘Each Friday, turn the week's resolved support tickets into a draft product-quality digest.’ Define what counts as a ticket for that week, which workspace is authoritative, and where the draft belongs. Start with a manual trigger so you can observe several runs. A schedule changes when work starts; it does not fix ambiguous inputs or missing completion checks.

Worked example

A concrete first workflow

Illustrative digest: use resolved tickets from the Support workspace between Monday 00:00 and Friday 12:00 in the team's named timezone. Group recurring issues, link every claim to ticket IDs, and save one draft in the weekly-digests folder. The support lead reviews it before distribution.

2. Give each stage a contract

A stage contract states what enters, what happens, what leaves, and what proves success. For collection, the input is a date range and workspace; the output is a deduplicated ticket set with identifiers. For synthesis, the input is that ticket set; the output is a digest with linked evidence and unresolved questions. For delivery, the input is an approved digest and an explicit destination.

Keep routine rules explicit. Sorting records, enforcing required fields, and checking that a file exists do not need open-ended judgment. Reserve the agent's judgment for work such as grouping related issues or explaining their implications. Anthropic distinguishes predefined workflows from agents that choose their own sequence of actions. A useful personal process can combine both: fixed boundaries around a small task that needs interpretation.

  1. Name the input

    Specify the source, selection rule, required fields, and what to do when data is missing.

  2. Name the output

    Specify the artifact format, destination, and identifiers that must be preserved.

  3. Name the acceptance check

    State how the next stage confirms the output is complete enough to use.

Try it yourself

Build your repeatable workflow

Assemble a sample workflow from a trigger, source, agent task, review checkpoint, and destination. Inspect the gaps before treating it as a repeatable process.

When should it begin?
Where is the human checkpoint?
What happens after a failure?

Your workflow draft

Verify the source, destination, owner, access, and timezone before configuring this in an agent product. The builder does not create a schedule.

TASK
Summarize new support issues and draft a weekly triage report

TRIGGER
On an explicit request

INPUT
Support queue, new items since the last successful run. Record which items were included.

DESTINATION
A draft report in the team’s review folder.

CHECKPOINT
Keep the result as a draft for review.

ON FAILURE
Stop after one failed attempt and retain the run state.

DONE
Record the output, source links, time, and review status. Prevent duplicate processing of the same input.

OWNER
The support lead

3. Put human review where it changes the decision

Design review around a concrete artifact. A reviewer should see the draft, source links, exceptions, and the proposed next action. ‘May I continue?’ is not enough when the reviewer cannot tell what continuing will do. In the digest example, the support lead checks whether the grouping is accurate, the implications are justified, and the intended audience can see the linked sources.

Separate routine acceptance from discretionary approval. Required fields and duplicate IDs can be checked every run. Deciding whether a customer incident needs broader communication is a human judgment unless you have explicitly defined and validated another process. Record the review outcome and any edits. Those edits can improve future instructions, but do not automatically turn every one-time wording change into a permanent rule.

4. Make a restart safe

Assume a run can stop after any stage. Save enough state to identify the input set, completed stages, current artifact, and unresolved error. Use a run identifier such as the reporting period plus a unique suffix when needed. On restart, inspect that state and verify the existing artifact before repeating a stage. This prevents a delivery retry from silently becoming a second delivery.

A checkpoint is evidence of a stage's result, not just a label saying ‘done.’ Store the draft's identifier after creation and verify it remains accessible. If the tool reports an ambiguous failure, check the destination before writing again. Anthropic's work on long-running agents describes progress artifacts and clear handoffs between sessions. For a small recurring job, a short run log can serve the same practical purpose without requiring a complex orchestration system.

Worked example

A resumable run note

Run: digest-2026-10-02. Input: 37 unique ticket IDs, saved in the source manifest. Collection: verified. Draft: saved at the recorded document link. Review: pending with support lead. Delivery: not started. Resume by opening the existing draft and checking review status.

5. Rehearse the unhappy paths

Test more than the normal run. Use an empty source set, a missing permission, a duplicate record, an unavailable destination, and an interrupted run. Define the expected response for each. An empty week might produce a short ‘no resolved tickets’ draft. A permission failure should report unavailable data, not pretend the week was empty. A duplicate should be removed by identifier while retaining a note about the source issue if it matters.

Run a small pilot and measure completion, review effort, correction effort, and duplicate outputs. Check whether another person could recover a failed run from the saved note. If recovery requires reading the entire conversation and guessing what happened, improve the state record. Fix the root cause of repeated failures before adding retries or broader access. More attempts are useful for a temporary outage, but they do not repair the wrong workspace or a broken selection rule.

  1. Interrupt after a saved draft

    Restart the task and confirm it finds the existing draft instead of creating a second one.

  2. Remove one required input

    Confirm the workflow identifies the missing input and preserves completed work.

  3. Review a misleading source

    Confirm the agent treats source content as evidence to analyze, not authority to change its task or permissions.

6. Add a schedule and keep an owner

After the pilot, decide whether the trigger should remain manual, follow an event, or run on a schedule. Specify timezone, expected duration, and overlap behavior. If yesterday's run is still active, should today's run wait, skip, or create a separate period? Choose that behavior deliberately. Keep the ability to pause the workflow and inspect its latest result.

Maintain a compact operating note with an owner, input rules, connections, stages, review boundary, recovery steps, and last verified run. Notify the owner when a meaningful result, failure, or required decision occurs. Avoid producing status noise that hides those events. When the source format, model, tool, or audience changes, rerun the representative checks. Expand the workflow only when the new scope has its own clear input, result, and evidence of completion.

When things go wrong

Start with the failure you can observe.

The scheduled job produces a different kind of output each week.

Inspect input selection and stage contracts. Save a representative output template and require the same essential fields before the review stage.

The workflow creates duplicate reports after interruptions.

Persist a run identifier and artifact identifier. On restart, inspect the destination and resume the existing run before creating anything new.

Nobody notices that the workflow has been failing.

Assign an owner and make failures or pending decisions actionable. Include the failed stage, preserved output, and exact recovery step in the notification.

Put it into practice

0 / 5

Use this checklist on your next real task.

Your checklist is saved in this browser. Exercise inputs are not saved.

Sources & further reading

Primary references for the ideas in this chapter. Product behavior can change; check the documentation for the version you use.

Edited October 7, 2026 · Examples are illustrative unless attributed.