- edited 3 files (+41 −12)
- ran 9 commands
- 4 checks passed
- refused 1 action
Checks
- npm test
- tsc --noEmit
- eslint src
- cargo test -p forge-duo
Jarvai builds coding agents that don't grade their own work. Protected tests decide when a task is finished, and every turn ends with a receipt built only from what actually ran.
Checks
No checks were run this turn — done is the model's claim.
Sentinel · 1 flag
tests claimed, no runner called
Origin
Jarvis took one sentence and handed it to a team of AI agents. A planner wrote goals, a lead turned them into contracts, parallel workers built in separate git worktrees, and an acceptance agent ruled on whether to ship. Each role could run on a different model provider.
In April we wrote down what shipped means: a person accepted the result, nobody rescued it by hand, and a check outside the agents confirmed the artifact. Then we ran our ledger against it.
23/53
runs received a ship verdict from Jarvis's own acceptance agent
0/53
runs met our written definition of shipped
Run ledger snapshot, 2026-04-19 · 6567545
We put the zero first in our own README, above the 23. Everything we have built since starts from that gap between what agents report and what holds up.
Forge
Forge is our second system: a local-first desktop app for coding agents, built on Electron, React and a Rust engine. You bring your own model keys. It is a working prototype and has not been released.
In duo mode, two builders from different model families take turns building and critiquing inside an isolated git worktree. The acceptance tests are restored from outside that worktree before every grading round. If an agent edits them, the edit is undone and the run is flagged.
dd8d0937acceptance_test_tampered
After every turn, a receipt lists what was edited, what ran, which checks passed and what Forge refused to do. Every line is computed from the turn's tool evidence. When no check ran, the receipt says so in plain words.
27195388src/ui/stream/doneReceipt.ts
Before an agent changes a file, Forge stores a content-addressed snapshot of it. The last fifty turns can be undone.
electron/file-history.ts
Self-audit
Forge once sealed a turn as completed after the stream ended without a result. It did the same for a tool call that never resolved. We counted both as fake dones and closed them.
4eb7c438a573c64b
Before testing duo mode against a single agent, we wrote down how the result would be judged. Both setups passed all 16 tasks, so we claim no accuracy difference.
0bad5a83
Jarvis keeps an append-only log of the numbers and citations its own agents got wrong.
docs/retraction_log.yaml
Forge's receipt copy bans the word verified, even in the negative. Checks passed is a fact. Verified would be a promise.
DoneReceiptCard.tsx
Build log
Jarvis v0.1.0. One sentence in; a planner, a lead, parallel workers and an acceptance agent take it from there.
1436d4f
The acceptance agent has to see the tests pass before it can ship.
b028637
It also has to launch the app it built and take a screenshot.
f6cbc93
Jarvis's own workers build part of its desktop app.
126a489
A strict definition of shipped. 0 of 53 runs qualify.
6567545
Forge's first commit.
e7bd5f5c
Forge's README opens with one line: your coding AI can't fake done.
372944a8
Night Shift: a scheduler for overnight runs and a summary card for the morning.
e69ec7fc
Duo engine: two model families, one isolated worktree.
dd8d0937
Receipt cards land. The packaged Windows build passes its release chain end to end.
271953880d7ed0db
Duo against a single agent, with the decision rule written first.
0bad5a83
Forge's own fake-done paths closed.
4eb7c438
Training open-weight models. The current approach learns from the model's own failed attempts, and a sandbox decides which fixes count.
Now
An agent harness can catch a fake done after the fact. We want models that produce fewer of them. Since September we have been training open-weight models. Our current approach teaches a model from its own failed attempts, with a sandbox deciding which fixes count.
Before we call any change a gain, we measure how far the score moves when nothing has changed at all. No results to report yet. We will post them when they clear that bar.
Contact
Every reference here points to a commit or a file in our private repositories. Pick one and we will walk you through it.