The blind AI model test

One prompt. Every model.
No names until you've judged.

We run one identical task through 13 AI models — 8 local open-weight models on our own Brain cluster and 5 cloud flagships. The answers appear anonymized as Model A, B, C … You read, compare and pick a favourite. Only then do you reveal who wrote what — together with blind rubric scores against a documented gold answer. No synthetic benchmarks, no cherry-picking: every raw answer is on the page, including the failures.

54rounds published, each with every raw answer
13models compared — 8 local open-weight, 5 cloud
8disciplines, from logic traps to worked physics problems

How a blind round works

STEP 01

One identical prompt

Every model in the field gets exactly the same task through the same API harness — same wording, one shot, no retries. Timeouts and errors count as a DNF and stay on the page.

STEP 02

Answers, anonymized

The raw answers are shuffled and labelled Model A, B, C … — no names, no scores, no latencies. You compare them the way a blind tasting works: on substance alone.

STEP 03

The reveal

When you're done judging, open the reveal: who wrote what, blind rubric scores against a gold answer, and response times. Did the model you picked win?

Latest blind rounds R31–R55

RoundTaskCategory
R55The double-slit experiment and wave-particle duality (conceptual)Physics
R54Time dilation: why do moving clocks run slow? (conceptual)Physics
R53Why is the sky blue? (conceptual)Physics
R52Fermi estimate: piano tuners in a 10-million-person cityPhysics
R51Heat required to warm water (specific heat capacity)Physics
R50Thin lens equation: image distance and magnificationPhysics
R49Ohm's law: resistance and power dissipationPhysics
R48Low-orbit satellite: speed and period (orbital estimate)Physics
R47Unit conversion: km/h to m/s and mphPhysics
R46Energy conservation on a frictionless inclinePhysics
R45Idiom-dense German to natural EnglishLanguage & Translation
R44Appointment extraction to strict JSONStructure & Data
R43Five machines (rate-reasoning trap)Knowledge & Logic
R42All but nine (word-problem trap)Knowledge & Logic
R41Prose → valid JSONStructure & Data
R40isPalindrome (ignore case & non-alphanumerics)Code
R3940-word product blurb with 3 required wordsCreative & Marketing
R38Register shift: formal → casual (German)Language & Translation
R37False friend: “control” EN→DELanguage & Translation
R36Continue the sequence (doubling gaps)Knowledge & Logic
R35Minutes in a (non-leap) year — show the mathKnowledge & Logic
R34Constrained taglines (word ban + length cap)Creative & Marketing
R33Idioms that break literal translation (EN→DE)Language & Translation
R32Four houses, four clues (constraint puzzle)Knowledge & Logic
R31Bat and ball (the classic reasoning trap)Knowledge & Logic

Earlier rounds R1–R30 are in the archive — they predate the blind-first layout and show model names directly.

Standings & scoreboards

We don't publish a single magic ranking number. Scored rounds carry blind rubric scores (0–10 against a gold answer, DNFs included); creative and style rounds stay deliberately unscored — read them yourself. The full picture lives on the overview page.

Eight disciplines

The field — 13 models

Eight local open-weight models running on our own hardware, five cloud flagships over their official APIs. Every profile links to all of that model's answers.

Why blind?

Model names carry expectations — a famous logo makes an answer look smarter than it is. Blind comparison removes that halo. Our method, its limits, and how scoring works are documented openly on the method page.