Research

What we study, and what we found

Published work comes with the data and code behind it, and says where its method fails.

  1. qed-bench: benchmarking metrics against task-appropriate baselines

    We trained metrics on four content-judgment tasks: holistic essay quality, SMS spam, AI-vs-human authorship, and LLM authorship attribution. Each one was compared to its task-appropriate baseline: trained human raters, gold labels, or an eight-model LLM-as-judge panel. Notebooks, model definitions, and per-judge artifacts at github.com/u22a8/qed-bench.

    read it

We also run The AI Oversight Study, on how organizations check and trust the automated systems they run, at oversight.study.