
AI judgment should know how likely it is to be right, and hold back where the evidence runs out.
We build models for AI judgment and evaluations. You state the criterion in plain words and set the accuracy. They answer where they can meet it, and say so where they can’t.
The calls
AI systems make judgment calls all day.
- Whether a reply meets the bar.
- Which queue a ticket belongs in.
- Whether an agent finished its task.
- Whether an answer is right.
- Whether a listing breaks the rules.
Each is a judgment: a criterion, a thing to hold against it, and a verdict. A careful person making these calls knows which ones are beyond them, and says so. That is what makes their verdicts worth acting on.
The two ways today
There are two ways to automate a judgment. Each costs something.
A trained classifier
It is consistent, and its accuracy is measured before it goes live. But every criterion needs its own labeled data, and someone to train it.
A generative model
It takes any criterion in plain words. But its verdict is sampled, not scored. Nothing in the answer says how likely it is to be right, and finding out takes the same labeled data the classifier needed. Ask again, and the verdict can change.
One knows its accuracy but learns each criterion from scratch. The other takes any criterion but cannot tell you its accuracy.
What we build · A8 (Touchstone), available for preview
A8 is a model that takes the criterion in plain words, and answers only where it can carry the accuracy you set.
You call it like any OpenAI-compatible chat model. The system message is the criterion. The user message is the thing to judge.
It answers above the bar you set, and abstains below it.
The bar is a number from 0 to 99, set on each call. At 95, at least 95 of every 100 verdicts A8 gives are correct, and a case the evidence can’t carry that far comes back as an abstention. The same request returns the same verdict.
It learns from every correction, and only for you.
Teaching a generative model a new criterion means fine-tuning it. Teaching A8 is one field on a request you already send: the verdict you expected, as expected.
| Fine-tuning a generative model | Teaching A8 |
|---|---|
| A labeled dataset, gathered and uploaded | The verdict you expected, on the request itself |
| A training run to start and watch | A fitting round in the background, a minute after your workspace goes quiet |
| A copy of the model to host | Nothing new to host |
Corrections are kept the moment they arrive, and answering never waits for fitting. What your workspace teaches trains only your workspace’s model. It never reaches another organization.
Automate the calls that clear your bar. The rest come back to you.
How we know
We test the approach against trained human raters, in the open.
In qed-bench, we scored essays in an encoder’s space, the approach A8 is built on, and compared the ranking with the one trained human raters gave. A panel of eight generative models judged the same essays.
- Our approach
- 0.815
- The strongest generative judge on the panel
- 0.813
Rank correlation with trained human raters, essay quality. At 1, the two rankings match exactly. The notebooks are public. read qed-bench
When a judgment is automated, its accuracy should be a number you set, not one you find out later.

Tell us where AI judgment falls short in your work.
We are asking the people who build, run and sign off on automated systems how they check them and decide when to trust them.
The AI Oversight Study is open until December 31, 2026. About five minutes, with questions for your role.
take the survey