a generative model
Fine-tuning
- A labeled dataset, collected first
- A training run to schedule
- A copy of the model to host
available for preview
Any criterion, in plain words. A verdict that clears the accuracy you set, or none at all.
the stone
Theophrastus described the test around 300 BC. An assayer drew the gold across a black stone and left a streak. Beside it he drew needles of known fineness, one streak each, and read the gold’s color against theirs.
The needle it matched gave the grade. Where no needle matched, there was no grade to give.
the same five parts
expected.min_accuracy, from 0 to 99, set on each call.the grade
Every call carries the accuracy its verdict must meet, from 0 to 99. At 95, at least 95 answered verdicts in a hundred are correct. Where the evidence can’t carry a case that far, A8 returns no verdict.
The same request returns the same verdict. No generative language model sits on the response path.
a call
The system message is the criterion: what you are judging.
The user message is the subject: the thing under judgment.
min_accuracy is the grade. At 95, at least 95 answered verdicts in a hundred are correct. Below it, A8 abstains.
expected is a correction. Your workspace’s model learns from it.
from openai import OpenAI
client = OpenAI(
base_url="https://api.u22a8.ai/eval/v1",
api_key="<your key>",
)
resp = client.chat.completions.create(
model="a8",
messages=[
{"role": "system", "content": "Does the reply answer it?"}, # 1
{"role": "user", "content": "Your parcel left today."}, # 2
],
extra_body={"min_accuracy": 95}, # 3
)
resp.choices[0].message.content # the verdict, or empty
# 4: the last verdict was wrong
client.chat.completions.create(
model="a8",
messages=[...],
extra_body={"expected": False},
)the needles
a generative model
A8
expectedA correction is banked the moment it arrives. About a minute after your workspace goes quiet, A8 fits a new version in the background, and the next verdicts come from it. Answering never waits for fitting.
Your corrections train your workspace’s model and no other. They never reach another organization.
the method, assayed
On essay quality, we compared scoring in an encoder’s space with trained human raters and with a panel of eight LLM judges. Its rank correlation with the raters was 0.815. The strongest judge on the panel reached 0.813.
read qed-benchState the criterion and the grade it must clear. Correct what A8 gets wrong, and your workspace’s model learns from it. A8 is available for preview.
To build A8 into your stack or a tool, write to .