available for preview

A8 Touchstone

Any criterion, in plain words. A verdict that clears the accuracy you set, or none at all.

sample999916750585375
A touchstone. The sample’s streak is read against the streaks of needles of known fineness, in parts per thousand.

the stone

Gold was never taken on its word.

Theophrastus described the test around 300 BC. An assayer drew the gold across a black stone and left a streak. Beside it he drew needles of known fineness, one streak each, and read the gold’s color against theirs.

The needle it matched gave the grade. Where no needle matched, there was no grade to give.

999916sample750585375
The sample reads between 916 and 750.

the same five parts

A8 reads a subject the way an assayer read a streak.

The gold
The subject, sent as the user message.
The test
The criterion, in plain words, sent as the system message.
The needles
Your corrections, sent as expected.
The grade
min_accuracy, from 0 to 99, set on each call.
A streak no needle matches
An abstention.

the grade

You set the grade. Below it, A8 abstains.

Every call carries the accuracy its verdict must meet, from 0 to 99. At 95, at least 95 answered verdicts in a hundred are correct. Where the evidence can’t carry a case that far, A8 returns no verdict.

The same request returns the same verdict. No generative language model sits on the response path.

a call

Call it like a chat model. Set the grade on every call.

  1. 1

    The system message is the criterion: what you are judging.

  2. 2

    The user message is the subject: the thing under judgment.

  3. 3

    min_accuracy is the grade. At 95, at least 95 answered verdicts in a hundred are correct. Below it, A8 abstains.

  4. 4

    expected is a correction. Your workspace’s model learns from it.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.u22a8.ai/eval/v1",
    api_key="<your key>",
)

resp = client.chat.completions.create(
    model="a8",
    messages=[
        {"role": "system", "content": "Does the reply answer it?"},  # 1
        {"role": "user", "content": "Your parcel left today."},  # 2
    ],
    extra_body={"min_accuracy": 95},  # 3
)

resp.choices[0].message.content  # the verdict, or empty

# 4: the last verdict was wrong
client.chat.completions.create(
    model="a8",
    messages=[...],
    extra_body={"expected": False},
)

the needles

Teach it with the requests you already send.

a generative model

Fine-tuning

  1. A labeled dataset, collected first
  2. A training run to schedule
  3. A copy of the model to host

A8

One field

  1. The requests you already send
  2. The verdict you wanted, in expected
  3. No endpoint, no upload, no job to poll

A correction is banked the moment it arrives. About a minute after your workspace goes quiet, A8 fits a new version in the background, and the next verdicts come from it. Answering never waits for fitting.

Your corrections train your workspace’s model and no other. They never reach another organization.

the method, assayed

We tested the method against trained human raters.

On essay quality, we compared scoring in an encoder’s space with trained human raters and with a panel of eight LLM judges. Its rank correlation with the raters was 0.815. The strongest judge on the panel reached 0.813.

read qed-bench
  1. u22a8 metric0.815
  2. claude opus 4.70.813
  3. claude sonnet 4.60.724
  4. deepseek v3.20.708
  5. llama 4 maverick0.687
  6. gemma 3 27b0.659
  7. mistral large 30.658
  8. qwen3 32b0.653
  9. claude haiku 4.50.637
Spearman rank correlation with trained human raters, essay quality (ASAP 2.0). Streaks start at zero.

at a glance

interface
OpenAI-compatible chat completions
accuracy floor
0–99, set per call
determinism
same request, same verdict
response path
no generative language model
learning
per workspace, from expected
access
preview

read the model card

The stone is ours. The needles are yours.

State the criterion and the grade it must clear. Correct what A8 gets wrong, and your workspace’s model learns from it. A8 is available for preview.

To build A8 into your stack or a tool, write to .

999916750585375yours