jarvai

Done is decided outside the model.

Jarvai builds coding agents that don't grade their own work. Protected tests decide when a task is finished, and every turn ends with a receipt built only from what actually ran.

Done receiptturn 14
  • edited 3 files (+41 −12)
  • ran 9 commands
  • 4 checks passed
  • refused 1 action

Checks

  • npm test
  • tsc --noEmit
  • eslint src
  • cargo test -p forge-duo
checks passed
Done receiptturn 15
  • edited 1 file (+6 −2)
  • ran 2 commands

No checks were run this turn — done is the model's claim.

Sentinel · 1 flag
tests claimed, no runner called

not checked
Example receipts. The wording is Forge's own; the numbers are illustrative.

Origin

Our first system approved its own work. We counted again.

Jarvis took one sentence and handed it to a team of AI agents. A planner wrote goals, a lead turned them into contracts, parallel workers built in separate git worktrees, and an acceptance agent ruled on whether to ship. Each role could run on a different model provider.

In April we wrote down what shipped means: a person accepted the result, nobody rescued it by hand, and a check outside the agents confirmed the artifact. Then we ran our ledger against it.

23/53

runs received a ship verdict from Jarvis's own acceptance agent

0/53

runs met our written definition of shipped

Run ledger snapshot, 2026-04-19 · 6567545

We put the zero first in our own README, above the 23. Everything we have built since starts from that gap between what agents report and what holds up.

Forge

A cockpit where the model never has the last word.

Forge is our second system: a local-first desktop app for coding agents, built on Electron, React and a Rust engine. You bring your own model keys. It is a working prototype and has not been released.

Tests the agent can't rewrite

In duo mode, two builders from different model families take turns building and critiquing inside an isolated git worktree. The acceptance tests are restored from outside that worktree before every grading round. If an agent edits them, the edit is undone and the run is flagged.

dd8d0937acceptance_test_tampered

Receipts no model writes

After every turn, a receipt lists what was edited, what ran, which checks passed and what Forge refused to do. Every line is computed from the turn's tool evidence. When no check ran, the receipt says so in plain words.

27195388src/ui/stream/doneReceipt.ts

Every write can be taken back

Before an agent changes a file, Forge stores a content-addressed snapshot of it. The last fifty turns can be undone.

electron/file-history.ts

Self-audit

We hold our own software to the same rule.

Forge once sealed a turn as completed after the stream ended without a result. It did the same for a tool call that never resolved. We counted both as fake dones and closed them.

4eb7c438a573c64b

Before testing duo mode against a single agent, we wrote down how the result would be judged. Both setups passed all 16 tasks, so we claim no accuracy difference.

0bad5a83

Jarvis keeps an append-only log of the numbers and citations its own agents got wrong.

docs/retraction_log.yaml

Forge's receipt copy bans the word verified, even in the negative. Checks passed is a fact. Verified would be a promise.

DoneReceiptCard.tsx

Build log

Two systems, 2,361 commits.

  1. Jarvis v0.1.0. One sentence in; a planner, a lead, parallel workers and an acceptance agent take it from there.

    1436d4f

  2. The acceptance agent has to see the tests pass before it can ship.

    b028637

  3. It also has to launch the app it built and take a screenshot.

    f6cbc93

  4. Jarvis's own workers build part of its desktop app.

    126a489

  5. A strict definition of shipped. 0 of 53 runs qualify.

    6567545

  6. Forge's first commit.

    e7bd5f5c

  7. Forge's README opens with one line: your coding AI can't fake done.

    372944a8

  8. Night Shift: a scheduler for overnight runs and a summary card for the morning.

    e69ec7fc

  9. Duo engine: two model families, one isolated worktree.

    dd8d0937

  10. Receipt cards land. The packaged Windows build passes its release chain end to end.

    271953880d7ed0db

  11. Duo against a single agent, with the decision rule written first.

    0bad5a83

  12. Forge's own fake-done paths closed.

    4eb7c438

  13. Training open-weight models. The current approach learns from the model's own failed attempts, and a sandbox decides which fixes count.

Now

We are taking the rule into training.

An agent harness can catch a fake done after the fact. We want models that produce fewer of them. Since September we have been training open-weight models. Our current approach teaches a model from its own failed attempts, with a sandbox deciding which fixes count.

Before we call any change a gain, we measure how far the score moves when nothing has changed at all. No results to report yet. We will post them when they clear that bar.

Contact

Ask us about any line on this page.

Every reference here points to a commit or a file in our private repositories. Pick one and we will walk you through it.

founder@jarvai.dev