NanoJudge Bench v1

How well each judge model ranks answers for a set of 30 subjective questions.

Loading…

How the benchmark works

Each model acts as a NanoJudge judge across 30 subjective questions. For every question the model runs thousands of pairwise comparisons, which Bradley-Terry statistics fold into a single ranked table. These questions have no external right answer, so the benchmark measures how closely each model’s ranking agrees with the consensus of the rest of the panel.