VeriCommand

Direct answer · evidence over confidence

How do you prove what an AI agent actually did?

Not by asking it. An agent sounds identical when it's right and when it's wrong — so the proof has to live in structure, not in tone. VeriCommand records every task, dispatch, signal, and return on a hash-chained record that is re-checked every time anyone reads it.

hash-chained
Every record binds to the one before it, and the whole chain is re-checked on every read. An edit in place shows up as a break; a record rewritten whole would line up too.
returns bind to dispatches
Who did what, under which instruction — on the record, not from memory. Self-graded findings count as assertions, not clearance.
verifies offline
Tamper-evident evidence packs check out offline with an open-source MIT tool: chain and manifest integrity, plus the seal when one is present. No account, no network.

The same discipline applies to the product itself: the download page publishes the installer's exact size and SHA-256. Check it yourself.

Why do AI coding agents need an audit trail?

Because absence of error is not evidence of progress. In one recorded incident on this product's own board, the operator spent a day believing four agents were building. Zero files had been written. Nothing errored — so nothing looked wrong. That day produced the rule the system now enforces structurally: progress is never reported from the absence of failure, and an agent that goes quiet looks like exactly what it is.

Using VeriCommand with Claude Code →

What makes a record tamper-evident rather than just a log?

A log can be edited after the fact; that's why logs make weak evidence. On a hash chain, every event — packet, dispatch, signal, return — binds to the record before it, and the whole chain is re-checked on every read. Modify or delete a record in place and the links after it stop lining up, so the next read shows the break. The limit, stated plainly: a record rewritten whole would line up too.

Can I verify it without trusting the vendor?

Yes, and that's deliberate. VeriCommand exports tamper-evident evidence packs that verify offline with an open-source, MIT-licensed tool (pip install flightdeck-verify): chain and manifest integrity, plus the seal when one is present — no account, no network call, no API that could lie to you. An audit trail you can only check by asking the vendor's server isn't an audit trail; it's a promise with extra steps.

The standalone verifier on PyPI →

What does an agent's "done" actually prove here?

A return binds to the specific dispatch that authorized it — so "who did what, under which instruction" is on the record. And the system is quietly rude to unverified confidence: findings an agent grades itself on count as assertions, not clearance. A return of nineteen self-marked criticals displays as "0 of 19 adjudicated" until someone other than the author confirms them. Confidence is not evidence, and the record refuses the exchange rate.

What does an audit trail not do?

It doesn't make the agent right — it makes the agent's claims checkable, which is a different and more honest promise. Verification is still a verb somebody performs. And for agents on your own machine, attribution is declared, not cryptographically authenticated: the chain proves the sequence and integrity of what was recorded, not the OS-level identity of the local process that wrote it. We'd rather state that limit than paper over it — a security page you can't falsify isn't one you should trust.

Evidence, not vibes.

The local desk is free. Put your agents on a hash-chained record that breaks when edited in place, counts self-graded findings as assertions, and checks out offline without trusting us.