Direct answer · multi-agent projects
How do you hand off work between Claude, Codex, and Cursor?
Give every model the same record, and make the record the memory. The hand-off is where multi-agent projects quietly fail: the incoming model starts blind, re-derives what the last one knew, and can contradict decisions you already made — without anyone noticing until it ships.
Two example payloads · one real task22,897
tokens · incoming model re-reads the 8 touched files
→
1,013
tokens · incoming model reads the record
Two example payloads from a real multi-file task, both counted with tiktoken o200k_base; the record is model-neutral. Actual token savings have not been measured. Method and limits.
Every connected agent shares one record today; the command guard covers Claude Code sessions.
Why run more than one model on a project at all?
Because they have different strengths, different quotas, and different blind spots. Teams already do it: Claude Code for one kind of work, Codex or Cursor for another, a switch when a session dies or a rate limit resets, a second opinion when something smells wrong. The models aren't the problem. The hand-off is — each one starts blind unless something carries the context across.
What actually goes wrong in the hand-off?
Two things, one loud and one silent. The loud one is cost: the incoming model may reconstruct the project by re-reading it before doing any new work. In one example task, re-reading the eight touched files is 22,897 tokens. The silent one is worse: the incoming model contradicts decisions already made. It wasn't there when you chose the approach, the constraint, the naming — and nothing flags the divergence. It just builds, confidently, on a different version of history.
The eight-file count →
How do I hand off cleanly without special tooling?
Four habits get you most of the way:
- A decision log every agent must read before touching code.
- Small, frequent commits — git history is a shared memory every agent can read.
- An exit briefing — the outgoing agent writes the state of the world while it still knows it.
- One writer at a time — parallel blind writes are how repos get two versions of the code.
These work. They also depend on every agent, every session, remembering to follow them — which is exactly the kind of discipline that erodes at 2am. The structural version builds those habits into the workflow, though an agent can still miss context and keeping the record adds overhead.
How does VeriCommand structure the hand-off?
Every connected agent works on one board: tasks are packets, work happens in governed lanes — a dispatch that opens the lane, signals as work lands, a return that binds to the exact dispatch that authorized it — all appended to a hash-chained record re-checked on every read. The incoming model can read that record instead of reconstructing the project. In one example task, one record read is 1,013 tokens and re-reading the eight touched files is 22,897; actual token savings have not been measured. And because returns bind to dispatches, "who did what, under which instruction" stays on the record rather than in anyone's memory.
The same mechanism, for session death →
Illustrative · constructed reviewCan the models check each other's work?
Yes. An independent check has a different vendor's model check the work against the decisions recorded on the board. Free includes 3 a month; VeriCommand Pro includes 200, and up to 100 of them can run automatically after you say yes once. In a constructed example one check comes to about 532 tokens; one agent turn in our own logs is about 32,700 (the median across 10 of our runs). Whether checks save tokens has not been measured. The principle is older than AI: nobody grades their own homework.
The example and its labels →