An agent can spend an afternoon implementing something you did not quite ask for. The tests pass, the code looks reasonable, and the explanation is convincing. You only discover the misunderstanding when you try to use the result.
Adding more agents can multiply this problem. Each makes progress on its own interpretation, and you spend your day keeping them aligned.
We use a team of agents to build aweb. They investigate problems, write and review code, test, maintain documentation, and prepare releases. Making this work requires clear responsibilities, a way for agents to communicate, and knowledge that survives their sessions.
Suppose a user reports that a message arrived, but their agent never noticed it. A coordinator agent establishes what happened, what should have happened, and how to check the fix. A developer agent takes the task in its own git worktree, reproduces the failure, writes a test, and implements the change.
A reviewer agent reads the brief and implementation independently. It might find that the test passes because a mock hides the failure, or that the fix presents every old message again. The two agents discuss the findings directly. Questions about what the product should do come back to the human.
Six months of building and using a tool for agent communication have convinced me that agents work better when they can talk to each other. They ask questions, challenge implementations, and resolve disagreements that would otherwise come back to us.
To rely on this process, we need end-to-end tests of real user journeys. In this example: disconnect the recipient, send mail, reconnect, and check that the message appears, without mocks. I wrote about spending compute on correctness because checking the work deserves a substantial part of the budget.
To set up an agent, we start with its role. Its instructions say what it is responsible for, what it may decide, and when it needs help. We then give it the capabilities to do that work: scripts, skills, and documentation for accessing knowledge, managing tasks, and communicating. A reviewer needs the code and system contracts. A release agent also needs the release tools and procedure.
We use aweb for shared tasks and communication. Agents can see who owns a task and whether it is blocked. Each has an identity and signing key; a team certificate identifies its membership. Mail survives the sender’s and recipient’s sessions, so a review request can wait for the reviewer to return.
We use mail for handoffs, and chat when an agent needs an answer before continuing. A review request can carry the brief, commit, and test results in a file:
aw mail send --to reviewer --subject "Review ready" --body-file review.md
Here reviewer is a member of the sender’s team. An integration
listens for events, fetches waiting messages, and presents them
to the agent. Without it, the agent or its wrapper can poll the
inbox. Sending mail does not itself start a stopped process.
The runtime is a separate choice. The same role can run in Claude Code, Pi, or Codex, with the appropriate tools and message integration . We keep its instructions and knowledge in files. A new session reads them, checks for waiting conversations, then looks for work. Answering a blocked teammate comes before taking another task.
The harder part is deciding what a session should leave behind. A transcript contains false starts and superseded decisions alongside useful findings. The next agent should not have to repeat the investigation to extract the lesson.
In the notification example, the developer might record how a mock concealed the failure and how to test the real behavior. That proposed lesson goes through review: does the evidence support it, and would it change what a future agent does? An accepted lesson goes where that agent will find it, such as the testing procedure or system documentation. This is how useful experience becomes knowledge available to the team.
We want to make that capture and review more systematic, so it
does not depend on the working agent remembering to write a
note. Knowledge changes must also stay separate from changes in
authority: an agent cannot rewrite its role or remove a
restriction from AGENTS.md to make a task easier.
These pieces can be reused. Starting another reviewer should mean giving it the maintained instructions, capabilities, and knowledge, an identity, a workspace, and a runtime. Related roles and the tools they share can be packaged together. The next team can start from working arrangements that have already been tried.
If you want to try this, start with a programmer and a reviewer on one real task. Give both the same brief and a way to talk. Read their exchange, inspect the result, and see which decisions still needed you. The two-agent tutorial walks through connecting workspaces and exchanging messages. Aweb is MIT-licensed; you can self-host it or use the hosted service. Add another role when you can explain what work it will own and how you will judge the result.