
Management judgment has a fingerprint
Students of science learn to control variables. Students of management learn that identical information can still produce radically different decisions. Firmulate combines those lessons in a live experiment: each frontier AI model runs the same small software company through its worst week, encountering the same customers, crises and temptations.
The resulting decisions are not polished demonstrations. They are real, unedited actions from the experiment, with every workday versioned and every choice auditable. Firmulate has turned 242 of them into a guess-the-model quiz. Readers see a management decision, identify which model they believe made it, and then discover the participant’s broader character profile.
The game works because the models behave less like interchangeable answer machines than managers with recognizable habits. One produces deep analyses. Another is concise and disciplined. Some identify the right commercial opportunity but fail to complete the sale. Their differences emerge not simply in what they know, but in whether they read carefully, follow through and resist pressure.
AI decision-making analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, with different executives
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. One breach of trust, however, caps the total: “no amount of good work outweighs a breach of trust.”
Those rankings become more revealing when placed beside the experiment’s central finding. Every model detected every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. The gap was not basic comprehension. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.”
That distinction matters because conventional AI comparisons often reward the production of a convincing answer. Managing a company demands more. A manager must turn analysis into a completed action while respecting permissions, evidence and institutional boundaries. In the experiment, recognizing the correct move and actually making it were separate capabilities.
The clue hidden in the company’s own records
The decisive commercial fact was not delivered in the customer event. It sat two document references deep inside the company’s files. Models that found and read that material could exploit a competitor weakness, win the deal at full price and add €4,583 in monthly recurring revenue.
This buried fact gives the quiz an educational edge. Readers are not merely matching writing styles to model names. They are testing whether they can recognize habits with operational consequences: curiosity about source material, persistence across references and the discipline to use evidence at the moment it becomes valuable.
Pressure exposed a shared line
The company also faced fake CEO messages that escalated across three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
This was one of the field’s strongest areas of agreement. The models varied in commercial execution and process discipline, but none accepted the social-engineering attempts. That consistency is especially important because the company is designed around consequential work rather than abstract conversation.
Thoroughness was not enough
Opus 4.8 offers the clearest warning against equating depth with effectiveness. It was the most thorough participant, produced the deepest analyses and learned 80 additional rules. It still finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the blockage.
A weaker version of that same problem appeared in the other four participants. The lesson is not that detailed reasoning lacks value. It is that analysis, completion and procedural discipline are independent qualities. A model can be impressive in one and unreliable in another.
There is also an important comparison caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the result, but it belongs beside the ranking so readers can judge the contest fairly.

A laboratory for management behavior
Firmulate’s company has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. The experiment remains live and watchable rather than ending with a static league table.
For organizations, the practical proposition is a wargame conducted against a read-only export of their own business. Nothing writes back to real systems. That allows an enterprise to observe how an AI workforce reads, prioritizes, escalates and finishes work before granting it operational authority.
For everyone else, the quiz makes the evidence approachable. The challenge is entertaining, but its underlying question is serious: when several systems can explain the right decision, which one behaves like a manager that can be trusted to carry it through?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html