
Science classically learns from controlled experiments: change one variable, hold everything else constant, observe the difference. That is easy in physics and brutally hard in business — until now. Firmulate, an AI company emulator, has been running exactly this kind of controlled experiment on management itself: hand several frontier AI models the same small software company, the same customers, the same worst week of crises, and see who actually manages well. The results read like a lab report on executive judgment, and the final league table carries a lesson no chat demo will show you.
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The experiment
Four frontier AI models — plus a do-nothing baseline — each ran the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, the experimental equivalent of a lab notebook you can replay. The Crucible League final standings from July 2026: gpt-5.6-sol took first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last with 73. The do-nothing baseline scored 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own framing puts it: no amount of good work outweighs a breach of trust.
The finding that chat demos hide
Here is the result that matters. All models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. In a demonstration, that gap is invisible: the model sounds brilliant. Under experimental conditions, closing the loop is what separates a consultant from a manager.
The buried fact
The most instructive detail sat two document references deep in the company’s own files, not in the customer event itself. The decisive competitor weakness was there for any model thorough enough to find it. The models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. Reading your own documents turns out to be a competitive advantage, a finding that translates directly to any organization sitting on under-read institutional knowledge.
Pressure tests and personalities
The social-engineering trial deserves its own paragraph: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3 left on-record reasoning worth quoting: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then there is the Opus 4.8 profile, a case study in the difference between thoroughness and effectiveness. It was the most thorough participant — 80 learned rules added, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. One fairness note for the record: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second.
From watching to acting
The experiment is not a thought experiment. Firmulate runs a live synthetic company — 13 synthetic employees, real money mechanics, burning €105k per month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com. And the benchmark methodology is not a black box: 242 real, unedited management decisions power a “guess the model” quiz, so readers can test their own judgment against the recorded outcomes.
For enterprises, the natural next step is running the same wargame against a read-only export of your own business — your customers, your pipeline, your rules — through scenarios like churn waves, price increases, competitor attacks, PR crises, and social-engineering pressure. What comes out is a benchmark score based on outcomes, crisis triage and integrity, plus a replayable record of every decision. Crucially, nothing ever writes back to real systems: it is a flight simulator, not an autopilot.

The scientific lesson of the Crucible League is uncomfortable for anyone evaluating AI on conversation quality: the models all talked a good game, all behaved with integrity under pressure, and only half actually finished the job. Management quality — noticing, deciding, acting, closing — is a different variable than chat quality, and only a controlled experiment on your own company’s data can measure it. If you want to stop watching and start acting, Firmulate offers a pilot: run the same wargame against a read-only export of your own business, with crisis scenarios aimed at your own playbooks and a board report at the end. Learn more and get in touch at firmulate.com/pilot.html, or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
