A8 model card

A8 (Touchstone) is an encoder-only evaluation model with a chat-model interface. It scores a subject against a criterion stated in plain words, and returns a verdict only where the evidence clears the accuracy the request sets.

Each section goes one level deeper than the last. For the short version, see A8.

At a glance

Versiona8-1
AccessPreview
InterfaceOpenAI-compatible chat completions
Base URLhttps://api.u22a8.ai/eval/v1
Accuracy floor0 to 99, set per call
ArchitectureEncoder-only. It scores text and generates none.
Generative language modelsNone on the response path
DeterminismSame request, same verdict
LearningCorrections sent as expected, fitted in the background
Data scopeYour organization’s model only

§1A call, and a correction

A8 answers the chat completions request you already send to a generative model. Only the base URL and the model name change.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.u22a8.ai/eval/v1",
    api_key="<your key>",
)

resp = client.chat.completions.create(
    model="a8",
    messages=[
        {"role": "system", "content": "Does the reply answer it?"},  # 1
        {"role": "user", "content": "Your parcel left today."},  # 2
    ],
    extra_body={"min_accuracy": 95},  # 3
)

resp.choices[0].message.content  # the verdict, or empty

# 4: the last verdict was wrong
client.chat.completions.create(
    model="a8",
    messages=[...],
    extra_body={"expected": False},
)
  1. The system message is the criterion: what you are judging.
  2. The last user message is the subject: the thing under judgment.
  3. min_accuracy is the floor. At 95, at least 95 answered verdicts in a hundred are correct. Where the evidence can’t carry a case that far, A8 abstains.
  4. expected is a correction. False says the verdict just returned was wrong. Send it with the same messages, and your organization’s model learns from it.

An optional json_schema names the fields the verdict comes back in.

The response

A standard chat.completion object. The verdict is the message content: a label, a number, or a JSON object matching the requested schema. This one answered a criterion with a schema asking for grounded.

{ "id": "eval-9f2c1d84ab30e5f7", "object": "chat.completion", "model": "a8-1-0@20260715t140322z", "choices": [{ "index": 0, "message": {"role": "assistant", "content": "{\"grounded\": true}"}, "finish_reason": "stop" }], "usage": {"prompt_tokens": 812, "completion_tokens": 0, "total_tokens": 812}, "system_fingerprint": "fit-1a2b3c4d5e6f7890.9e2a1b3c" }

model names the exact horizon that answered. Send it back to reproduce the verdict. Every field is in the reference.

§2Scored, not sampled

A generative model asked to judge writes its verdict the way it writes anything else: it samples it. Nothing in the answer says how likely it is to be right, finding out takes labeled data, and asking again can change it.

A8 scores the subject against what it has learned about the criterion, and answers only where that evidence clears the accuracy you set. No text is generated on the way, so the same request returns the same verdict.

When you need text back, such as an explanation, a rewrite or a summary, a generative model is the right tool. Use it to write, and A8 to check.

In a regression suite, a changed verdict means the input changed or the model moved.

Two things to know aboutThe horizon you pin freezes a verdict. The system_fingerprint in the response tells you something moved.

§3How it works

Encoders place the criterion and the subject in one space. For each criterion it has learned, the eval model holds a fitted direction in that space and the evidence around it. The verdict is where the subject falls along that direction.

Fitting runs in the background, and answering does not wait for it. Each fit mints a new horizon, and the response names the one that answered.

§4Abstention and the accuracy floor

min_accuracy is a promise about answered verdicts: an integer from 0 to 99, default 90. At 95, at least 95 in a hundred are correct. The cases the evidence does not cover leave the answered set instead of entering it as guesses, which is what keeps the rate true.

Every answer carries the level it is promised at: the level asked for, or the nearest earned one above it. 0 asks for no promise. A subject unlike anything the model learned from abstains at any level above that.

It is a floor the request sets, not a belief the model reports. Nothing in the response says a particular answer is 95% likely to be right. The number is the accuracy the evidence had to show before the model returned a verdict.

An abstention on the wire

The request still succeeds with 200. The verdict comes back null, or empty where the request carried no schema, and a top-level abstention names the check the evidence failed.

{ "choices": [{ "message": {"content": "{\"meets\": null}"}, ... }], "abstention": {"reason": "margin 0.31 below the 90% cut 0.44"} }

The checks, and what clears each, are in abstention.

§5Learning from corrections

A correction is an ordinary request that carries expected. There is no dataset to upload, no training job to schedule or poll, and no tuned model to host.

The correction is banked immediately. A fit follows once your organization has been quiet for about a minute, so a hundred corrections sent in a loop become one fit. When it completes, a new horizon starts serving, and the one before it stays pinnable.

What your organization teaches trains your organization’s model. It never reaches another organization.

The three forms of expected
ValueMeaning
A verdict objectTrain on this answer for this subject.
trueThe verdict just returned was right. Reinforce it.
falseThe verdict just returned was wrong. Record “not this answer”.

The shorthands do nothing after an abstention: there is no verdict to affirm or deny. Corrections, withdrawals and scope are in fine-tuning.

§6Evidence

qed-bench measured the approach A8 is built on: scoring in an encoder’s space. On essay quality (ASAP 2.0), its rank correlation with trained human raters was 0.815. The strongest LLM judge on the panel, claude opus 4.7, reached 0.813.

Scopeqed-bench measured the scoring approach on public datasets, not on your criteria. For yours, the floor you set on each call is the promise.

§7What to expect

It abstainsWhen the evidence falls short of the floor you set, the response says so.
Pinned versions stay fixedPin a horizon, and its verdicts stay the same. Every response names the model that answered.
Corrections train itSend the verdict you expected, and your organization’s model learns it. Withdraw the example, and the next fit leaves it out.
Your data stays yoursWhat your organization sends trains only your organization’s model.
It declines unfamiliar subjectsOn a subject unlike anything it learned from, it abstains.
It is deterministicThe same request against the same horizon returns the same response, byte for byte.

§8Compared with the other two ways

A judgment is automated today with a trained classifier or a generative model acting as judge. This is how the three differ.

Trained classifierLLM judgeA8
The criterionFixed by the training labels. One model per criterion.Plain words in the prompt.Plain words in the system message.
OutputA label or a score.Sampled text, parsed into a verdict.A scored verdict, or an abstention.
AccuracyMeasured on held-out labeled data before use.Unknown until you label a set and measure it.Set per call with min_accuracy. Below it, the answer is withheld.
Same input, same verdictYes.Not guaranteed. Sampling can change it.Yes, against the same horizon, byte for byte.
Labeled data to startA labeled dataset per criterion.None to answer. A labeled set to know its accuracy.None. Where the evidence is thin, it abstains.
Adapting to a correctionRelabel, retrain and redeploy.Rewrite the prompt, or fine-tune: a dataset, a training run, and hosting for the tuned model.One expected field on an ordinary request. A background fit follows about a minute after the last correction.
What a correction reachesThe model you retrain.The prompt or the tuned model you deploy.Your organization’s model only.
Generative model on the response pathNo.Yes. It is the judge.No. Nothing writes the verdict: completion_tokens is always 0.