Research
What we study, and what we found
Published work comes with the data and code behind it, and says where its method fails.
qed-bench: benchmarking metrics against task-appropriate baselines
We trained metrics on four content-judgment tasks: holistic essay quality, SMS spam, AI-vs-human authorship, and LLM authorship attribution. Each one was compared to its task-appropriate baseline: trained human raters, gold labels, or an eight-model LLM-as-judge panel. Notebooks, model definitions, and per-judge artifacts at
read itgithub.com/u22a8/qed-bench.
We also run The AI Oversight Study, on how organizations check and trust the automated systems they run, at oversight.study.