The Pop Quiz That Beat the Demo: How a Live Company Became the Ultimate AI Exam
Five frontier AIs ran the same company through its worst week. A Moonshot newcomer nearly beat them all — and exposed why chat demos can’t predict management skill.
Why the Worst AI Manager Still Gets a 26: A Lesson in Honest Grading
A do-nothing AI manager scores 26, not 0 — and no model ever hits 100. Inside the grading philosophy of a benchmark that caps the score on a single breach of trust.
The Exam AIs Didn’t Know They Were Taking: When a Buried Footnote Decided a €55,000 Deal
Four frontier AIs ran the same company through its worst week. All passed the honesty tests — but only those who read a fact buried two references deep closed the €55,000 deal.
Four frontier AIs ran the same company through its worst week. All passed the honesty test — only two signed the €55k deal. Management quality beats chat quality.