VeriCommand

Benchmark · two example payloads

Does keeping the work on one record actually save tokens?

We have not measured that. What this page shows is narrower: two example payloads from one real multi-file task, one for resuming without VeriCommand’s record and one for resuming with it, each counted with the same tokenizer.

Actual token savings have not been measured.

In this example, re-reading the eight files the task touched comes to 22,897 tokens, and one read of VeriCommand’s record for the same task comes to 1,013 tokens. Both were counted with the tiktoken o200k_base tokenizer, a proxy for Claude’s own. Reading one compact record can cost fewer tokens than re-reading the files it describes, as it does here. Whether it does in your work depends on how your sessions would otherwise resume.

Example payloads · token counts, not savings

8
files the example task touched
22,897
tokens to re-read those eight files
1,013
tokens for one read of the record
427
tokens of governance calls in one example session

01What was counted

This page looks at one part of the token question: what it can cost to re-establish working state at a context boundary (a session resume, a hand-off to another model, or a restart after the context window compacts), and what the record adds while you work.

The task is real: RatePilot’s team-seats feature, which touched eight files across billing, server, data access, templates and tests. Every payload below was counted with tiktoken o200k_base, a proxy for Claude’s own tokenizer, applied the same way to each.

  • Payload one: the eight files. One way to resume without a record is for the agent to re-read the files it touched, to work out what it decided and what is left. This payload is the full text of those eight files.
  • Payload two: one record read. With the record, the agent can instead make one read_truth call and get a compact state block from the hash-chained record. This payload is that block for this task, captured from a real record.
  • The governance calls. Keeping the record also uses tokens while you work: creating the task, dispatching it, sending progress signals and returning it. A representative set for one session on this task (create, dispatch, five progress signals, return) comes to 427 tokens.

02The numbers

Payload one: the files an agent would re-read to resume without the record.

File re-read to reconstruct stateTokens
server.py11,276
billing.py2,427
subscribe.html2,368
team.html1,731
test_team_seats.py1,670
supabase_rest.py1,530
test_trial_metering.py1,034
test_billing_scope.py861
Payload one: eight files22,897
Payload two: one record read1,013
Governance calls in one example session: create + dispatch + 5 progress signals + return427
Without record
22,897 — re-read 8 files
With record
1,013

These are token counts of two example payloads, not a measured saving. No live run compared resuming with and without the record.

03What the run logs show

VeriCommand’s own flightdeck_burn tool reads actual run logs. Across 29 real runs it measured 5.65M tokens (13 with reported usage; the rest logged honestly as unknown, never zero). The most expensive single run used 858,432 tokens across 43 turns. The tool’s diagnosis of that run: it re-ran an unnarrowed full test suite seven times, every full run returned its whole output into context, and context is re-read each turn, so the cost compounded. The same mechanism can apply after a resume: whatever an agent re-reads to rebuild its state then sits in context and is re-read on every turn. A compact record can keep one kind of bulky context out at a session boundary; it does nothing about output an agent pulls in during a session.

04Honest limits

  • No saving is measured here. These are token counts of two example payloads from one task. No live end-to-end run compared resuming with and without the record, so actual token savings have not been measured.
  • Proxy tokenizer. o200k_base is not Claude’s exact tokenizer, so Claude’s own counts of these payloads would differ.
  • Payload one is one way to resume, not the only way. It counts only the text of the eight files. A real resume may also search, list directories and read git history, which would add tokens. It may also re-read fewer files, or start from a compaction summary or the agent’s own notes, which would cost fewer. When the next session would otherwise resume from a short summary, the record read plus its overhead can cost more tokens than the summary does.
  • One session gets no benefit. In one continuous session there is nothing to resume, so there is nothing for the record read to replace, and the governance calls are overhead.
  • One task. The counts come from one task in one codebase. Other tasks would give other numbers.
  • What the record does not do. The record gives an agent the earlier decisions to consult; it does not make the agent follow them. Pro’s independent check is a separate question, covered in part two.

05What this supports

Not “VeriCommand saves tokens”; that has not been measured. It supports something smaller: in this example, one read of the record is 1,013 tokens, and re-reading the eight files it describes is 22,897. Reading one compact record at a session or model boundary can cost fewer tokens than rebuilding state from the files. Work that stays in one continuous session gets no benefit from it.

Method documented · scripts on request · counted 2026-09-02 · tiktoken o200k_base · real RatePilot team-seats task · one record read captured from a real record · flightdeck_burn run logs

Part two · Pro

What does one independent check cost in tokens?

A different question from everything above. Pro’s independent check has a model from a different vendor read your recorded decisions and the agent’s result, and look for conflicts. A check spends tokens to run. Whether it saves more than it spends depends on whether it finds a real conflict, and on how much work that conflict would have cost. We have not measured that.

Illustrative · constructed review, estimated turn cost

~32.7K
one agent turn (estimate: median of 10 of our runs)
532
one focused review (a constructed example)
  • What is and isn’t measured. The review cost is a constructed example: the token count of a sample review we wrote, not a measured review. The per-turn cost is an estimate: the median across 10 of our own agent runs in flightdeck_burn logs. How many turns a check might prevent has not been measured, so no savings figure is claimed.
  • It depends on a catch. A check can only prevent rework when it finds a real conflict. A check that finds nothing still spends its tokens.