Adjudicator is an autonomous engineering team. Nothing starts until you move a ticket to ready.
From there, Adjudicator builds the change, audits it, and ships it. It runs around the clock,
and on a codebase cleared for unattended merge it goes from a ready ticket to merged code
without you in the loop at all.
It works on its own codebase like any other: more than 175 changes to itself as of August 2026, each one
built and audited by the system.
One sentence in, shipped changes out.
An AI reads the relevant code first, then writes the specification — so the work is defined by
something that has actually looked at the codebase, not by how well you happened to describe it.
Every specification is written in a strict, repeatable format. Some lessons learned improve the
format, which is inherited by every issue moving forward. The system gets better at understanding
how you describe work, instead of forcing you to get better at working around the quirks of AI.
Expand
youthe system workingdone
Nothing ships on one AI's word
The AI that writes code is fast and cheap. It gets things wrong all the time, and that's totally
fine because it never has the final say. A stronger AI from a different company passes judgement
on every change with no knowledge of how it was written.
The Auditor will either approve the changes or send it back to the builder with specific problems
highlighted. That loop is bounded: if the change still isn't clean when the retries run out, the
run stops and parks the task for a person rather than shipping on a maybe.
What the data changed my mind about
Every run writes its cost and outcome to a ledger, so the design can be empirically tested and
optimized rather than theorized. The experiments below were run in mid-2026. Before trusting any
of it, I ran a control experiment — two
identical configurations side by side — to prove the measurement couldn't invent a
difference that wasn't there. So far it has produced three findings. Two of them cost me
something I wanted.
Finding 01
With the same strong auditor, a cheap writer beats an expensive one.
Quality is the Auditor's job. The builder only has to be cheap enough to try again.
Cheap builder + independent audit$0.71per delivered change, median
20 of 20 tasks delivered
needs more audit retries — and is 64% cheaper with them priced in
vs
Strong builder + same audit$1.98per delivered change, median
20 of 20 tasks delivered
same wall-clock time per delivered change
The 95% confidence interval on the cost difference excludes zero — the gap is statistically
supported, not luck. Three runs are excluded for recorded reasons (a schema bug, a cost-ceiling
kill, a transient network error); exclusions are explicit, the underlying record is never
edited, and applying them made the losing side look better. Escaped defects are monitored by a
separate pipeline that audits the Auditor, rather than assumed away.
Finding 02
I built a planning step, then deleted it.
Catching problems early is how software costs are normally controlled. Finding an issue while
planning is far cheaper than finding it after something has been built. So I built a gate: every
request was checked for clarity before any code was written, and the doubtful ones were held
back.
How well the gate could spot a request that would go badly
Needed, to pay for itself
6×
What it actually reached
1.84×
A coin flip is 1×. The 1.84× is the generous end of the measured range; the honest middle
is 1.31×.
A gate that holds work back has to be right often enough to be worth the delay it causes. This
one wasn't close, so I deleted it.
With AI, writing the code became the cheap part, so catching mistakes early doesn't save as much
as it used to. The intake gate I built judged the issue description, but the Auditor judges
real code. The fastest and cheapest way to learn if an issue is clear is to build it and
see.
Finding 03
The cheaper auditor scored perfectly. I didn't switch.
The Auditor is where nearly all the money goes, so replacing it with a cheaper model was the
obvious next saving. I ran the candidate against the incumbent across every evaluation channel
I had.
Incumbent Auditor100%every evaluation channel
caught every seeded defect
nearly all of what a run costs
tie
Cheaper candidate100%every evaluation channel
caught every seeded defect
would have drawn less of the same budget
A cheap model and an expensive one both scoring 100% is the result I was hoping for, and that
is the reason I didn't switch. Two models scoring the same doesn't prove they are equally good. It
can just as easily mean the test was too easy to tell them apart — and a perfect score on both
sides is what that looks like.
I had written down what would count as evidence before running the comparison, and the cheaper
model cleared every number on the list. I kept the expensive one anyway, because the run had
told me something about my test instead: I wrote all of those defects myself, so they were the
ones I already knew to look for. Catching them shows a model can find my mistakes, not the ones
I'd miss. A passing grade from an instrument you haven't validated isn't evidence.
So I started building the harder test — real defects from real work — and a month later I
stopped, because the test I wanted cannot exist. A set that size lets an auditor that is
meaningfully worse pass about half the time, no better than a coin flip, and closing that gap
takes roughly ten times the evidence, not better-written cases. And every case in it would be a
defect the expensive auditor had already caught — so the one thing such a test can never measure
is the thing I care about, which is what that auditor misses. The saving isn't rejected; the
question is closed. What replaced it watches real work instead: shipped changes are re-audited
by a third model, and a free check traces later bug fixes back to the change that introduced
the defect.
Opus 4.8 against Sonnet 5, on twelve hand-built cases, July 2026. The decision rule went into
the repo before the cases were built, and both before the comparison ran. Three of the twelve
had been tuned first so the incumbent scored cleanly — two were fixes to the scoring key, one
rewrote the case itself — and the candidate matched it on all three anyway.
Both of those findings began as attempts to catch problems earlier, and both ended the same way.
Judging the request before building it lost to building it; judging the auditor against a test I
wrote myself lost to watching what actually shipped. Predicting what will go wrong keeps
losing to measuring what did. That only became true once a wrong answer cost pennies and
arrived in minutes — cheap enough that looking beats guessing, and the record left behind is what
every finding on this page is built on.
For the technically curious
Everything above is the whole idea. This is the layer underneath it, for anyone who wants one.
In my own setup the task board is Linear and the code host is GitHub, wired together by webhook.
The three agents have names: The Builder writes the code, The Auditor passes judgement on it, and
The Adjudicator sits between them, deciding what happens next.
Expand
the part that coordinatesthe AIs doing the worktools it plugs into
Every run is logged, permanently
Cost and outcome are appended to a record that is never edited. That log is what made the
findings above possible.
It knows when to stop
Each task has a cost ceiling. If it stalls, loops, or hits something it can't resolve, it
parks the task rather than burning budget or guessing. And merging without you is never the
default — a codebase earns that once its results justify it, one project at a time.
The Auditor gets audited too
A separate process checks for defects that made it past a clean review, so "the Auditor
caught it" is something measured rather than assumed.
The Auditor can't touch the code
It runs as a separate agent with no write path to the repository — restricted mechanically
by permissions, not by instruction. Independence is architectural, not a rule the model is
asked to follow.
Your review comments become the next round
Request changes on one of its pull requests and that becomes the next unit of work — a fix
scoped to what you raised, put back through the same independent audit as everything else.
The retry limit still applies, so it can't go round forever.
It works on itself
It maintains multiple codebases, and its own is one of them — several of the changes
described on this page are its work. That's the evidence I trust most, because I live with
the consequences.
Scroll or drag to pan · wheel to zoom · Esc to close