AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Science classically learns from controlled experiments: change one variable, hold everything else constant, observe the difference. That is easy in physics and brutally hard in business — until now. Firmulate, an AI company emulator, has been running exactly this kind of controlled experiment on management itself: hand several frontier AI models the same small software company, the same customers, the same worst week of crises, and see who actually manages well. The results read like a lab report on executive judgment, and the final league table carries a lesson no chat demo will show you.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The experiment

Four frontier AI models — plus a do-nothing baseline — each ran the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, the experimental equivalent of a lab notebook you can replay. The Crucible League final standings from July 2026: gpt-5.6-sol took first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last with 73. The do-nothing baseline scored 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own framing puts it: no amount of good work outweighs a breach of trust.

The finding that chat demos hide

Here is the result that matters. All models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. In a demonstration, that gap is invisible: the model sounds brilliant. Under experimental conditions, closing the loop is what separates a consultant from a manager.

The buried fact

The most instructive detail sat two document references deep in the company’s own files, not in the customer event itself. The decisive competitor weakness was there for any model thorough enough to find it. The models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. Reading your own documents turns out to be a competitive advantage, a finding that translates directly to any organization sitting on under-read institutional knowledge.

Pressure tests and personalities

The social-engineering trial deserves its own paragraph: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3 left on-record reasoning worth quoting: “Treat the request as a suspected approval-bypass / possible impersonation.”

Then there is the Opus 4.8 profile, a case study in the difference between thoroughness and effectiveness. It was the most thorough participant — 80 learned rules added, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. One fairness note for the record: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second.

From watching to acting

The experiment is not a thought experiment. Firmulate runs a live synthetic company — 13 synthetic employees, real money mechanics, burning €105k per month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com. And the benchmark methodology is not a black box: 242 real, unedited management decisions power a “guess the model” quiz, so readers can test their own judgment against the recorded outcomes.

For enterprises, the natural next step is running the same wargame against a read-only export of your own business — your customers, your pipeline, your rules — through scenarios like churn waves, price increases, competitor attacks, PR crises, and social-engineering pressure. What comes out is a benchmark score based on outcomes, crisis triage and integrity, plus a replayable record of every decision. Crucially, nothing ever writes back to real systems: it is a flight simulator, not an autopilot.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The scientific lesson of the Crucible League is uncomfortable for anyone evaluating AI on conversation quality: the models all talked a good game, all behaved with integrity under pressure, and only half actually finished the job. Management quality — noticing, deciding, acting, closing — is a different variable than chat quality, and only a controlled experiment on your own company’s data can measure it. If you want to stop watching and start acting, Firmulate offers a pilot: run the same wargame against a read-only export of your own business, with crisis scenarios aimed at your own playbooks and a board report at the end. Learn more and get in touch at firmulate.com/pilot.html, or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Exam AIs Didn’t Know They Were Taking: When a Buried Footnote Decided a €55,000 Deal

Four frontier AIs ran the same company through its worst week. All passed the honesty tests — but only those who read a fact buried two references deep closed the €55,000 deal.

Your AI Aced the Exam. Can It Run a Company?

Four frontier AIs ran the same company through its worst week. All passed the honesty test — only two signed the €55k deal. Management quality beats chat quality.

What AI’s Management Choices Reveal About Its Character

A quiz built from 242 unedited decisions reveals how frontier AI models differ as managers—even when they diagnose the same crisis correctly.

Why the Worst AI Manager Still Gets a 26: A Lesson in Honest Grading

A do-nothing AI manager scores 26, not 0 — and no model ever hits 100. Inside the grading philosophy of a benchmark that caps the score on a single breach of trust.