A small sphere resting at the edge of a ridge, casting one long shadow.

AI judgment should know how likely it is to be right, and hold back where the evidence runs out.

We build models for AI judgment and evaluations. You state the criterion in plain words and set the accuracy. They answer where they can meet it, and say so where they can’t.

The calls

AI systems make judgment calls all day.

  • Whether a reply meets the bar.
  • Which queue a ticket belongs in.
  • Whether an agent finished its task.
  • Whether an answer is right.
  • Whether a listing breaks the rules.

Each is a judgment: a criterion, a thing to hold against it, and a verdict. A careful person making these calls knows which ones are beyond them, and says so. That is what makes their verdicts worth acting on.

The two ways today

There are two ways to automate a judgment. Each costs something.

A trained classifier

It is consistent, and its accuracy is measured before it goes live. But every criterion needs its own labeled data, and someone to train it.

A generative model

It takes any criterion in plain words. But its verdict is sampled, not scored. Nothing in the answer says how likely it is to be right, and finding out takes the same labeled data the classifier needed. Ask again, and the verdict can change.

One knows its accuracy but learns each criterion from scratch. The other takes any criterion but cannot tell you its accuracy.

What we build · A8 (Touchstone), available for preview

A8 is a model that takes the criterion in plain words, and answers only where it can carry the accuracy you set.

You call it like any OpenAI-compatible chat model. The system message is the criterion. The user message is the thing to judge.

It answers above the bar you set, and abstains below it.

The bar is a number from 0 to 99, set on each call. At 95, at least 95 of every 100 verdicts A8 gives are correct, and a case the evidence can’t carry that far comes back as an abstention. The same request returns the same verdict.

  1. Does the reply answer the customer’s question?

    “Your parcel left our warehouse this morning.”

    yes

  2. Which queue: billing, shipping or account?

    “I was charged twice for one order.”

    billing

  3. Is the answer free of hedging?

    “It may work, but I could be wrong.”

    no

  4. Does the message ask for a refund?

    “Can I swap this for a larger size?”

    abstainsevidence reaches 88

  5. Does the summary keep every figure in the source?

    “Revenue rose in the third quarter.”

    abstainsevidence reaches 82

  6. Is the review about the product or the delivery?

    “Arrived late, but it works well.”

    abstainsevidence reaches 64

answered 3 of 6 abstained 3Illustration. The rows and their evidence are invented to show the bar at work. They are not measured results.

It learns from every correction, and only for you.

Teaching a generative model a new criterion means fine-tuning it. Teaching A8 is one field on a request you already send: the verdict you expected, as expected.

Fine-tuning a generative modelTeaching A8
A labeled dataset, gathered and uploadedThe verdict you expected, on the request itself
A training run to start and watchA fitting round in the background, a minute after your workspace goes quiet
A copy of the model to hostNothing new to host

Corrections are kept the moment they arrive, and answering never waits for fitting. What your workspace teaches trains only your workspace’s model. It never reaches another organization.

Automate the calls that clear your bar. The rest come back to you.

How we know

We test the approach against trained human raters, in the open.

In qed-bench, we scored essays in an encoder’s space, the approach A8 is built on, and compared the ranking with the one trained human raters gave. A panel of eight generative models judged the same essays.

Our approach
0.815
The strongest generative judge on the panel
0.813

Rank correlation with trained human raters, essay quality. At 1, the two rankings match exactly. The notebooks are public. read qed-bench

When a judgment is automated, its accuracy should be a number you set, not one you find out later.

oversight.study

Tell us where AI judgment falls short in your work.

We are asking the people who build, run and sign off on automated systems how they check them and decide when to trust them.

The AI Oversight Study is open until December 31, 2026. About five minutes, with questions for your role.

take the survey