One prompt. Every model.
No names until you've judged.
We run one identical task through 13 AI models — 8 local open-weight models on our own Brain cluster and 5 cloud flagships. The answers appear anonymized as Model A, B, C … You read, compare and pick a favourite. Only then do you reveal who wrote what — together with blind rubric scores against a documented gold answer. No synthetic benchmarks, no cherry-picking: every raw answer is on the page, including the failures.
How a blind round works
One identical prompt
Every model in the field gets exactly the same task through the same API harness — same wording, one shot, no retries. Timeouts and errors count as a DNF and stay on the page.
Answers, anonymized
The raw answers are shuffled and labelled Model A, B, C … — no names, no scores, no latencies. You compare them the way a blind tasting works: on substance alone.
The reveal
When you're done judging, open the reveal: who wrote what, blind rubric scores against a gold answer, and response times. Did the model you picked win?
Start here — a classic reasoning trap
R31 · Bat and ball
A bat and a ball cost 1.10 together, the bat costs 1.00 more than the ball — the question almost everyone answers wrong on instinct. Eight models took it blind. Can you tell the careful reasoners from the pattern-matchers without seeing the names?
Latest blind rounds R31–R55
Earlier rounds R1–R30 are in the archive — they predate the blind-first layout and show model names directly.
Standings & scoreboards
We don't publish a single magic ranking number. Scored rounds carry blind rubric scores (0–10 against a gold answer, DNFs included); creative and style rounds stay deliberately unscored — read them yourself. The full picture lives on the overview page.
Eight disciplines
Knowledge & Logic
Factual knowledge without hallucinations and clean step-by-step reasoning on classic logic riddles and traps.
Physics
Worked physics with checkable numbers: mechanics, orbits, circuits, optics, heat, Fermi estimates — plus conceptual questions on relativity and quanta.
Code
Programming in a small space: correct one-liners, CSS that runs, and algorithm races rendered in a real browser.
Language & Translation
Language feel in German and English: translation, idioms, false friends, register and audience fit.
Structure & Data
Turning messy text into exact formats: schema-true JSON, classification, table understanding, format repair.
Creative & Marketing
Copy under constraints: slogans, micro-copy, product names and image prompts with market fit.
Physics Lab
Compact canvas physics sims — gravity, chaos, orbits — written one-shot and verified in a headless browser.
The field — 13 models
Eight local open-weight models running on our own hardware, five cloud flagships over their official APIs. Every profile links to all of that model's answers.
Why blind?
Model names carry expectations — a famous logo makes an answer look smarter than it is. Blind comparison removes that halo. Our method, its limits, and how scoring works are documented openly on the method page.