
Every educator knows the difference between a student who studied and a student who skimmed. You can hide one hard fact deep in the reading list, and the exam will quietly sort the class into those who did the reading and those who didn’t. No trick question required — just a fact placed two references deep.
In July 2026, an unusual experiment applied exactly that principle to frontier AI models. Four of the most capable systems on the market were each handed the same job: run a small software company through its worst week. All of them passed the obvious tests — every crisis spotted, every manipulation attempt refused. But buried in the company’s own files, two document references deep, sat a competitor weakness that would decide a €55,000 deal. Only the models that actually read the file closed it. The others left the money on the table.
The lesson is one teachers have given for centuries: you can’t fake having done the reading. Now it turns out that’s a measurable, purchase-deciding property of AI agents too.
The experiment
Firmulate, a public project that runs AI models as complete simulated companies, staged what it calls the Crucible league: each frontier model ran the same small software firm through identical customers, identical crises, and identical temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so the whole thing can be replayed rather than taken on faith.
The final standings: gpt-5.6-sol finished first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last with 73. A do-nothing baseline — an agent that simply sat on its hands — scored 26, a reminder that partial progress counts, but that a single breach of trust caps the total. As the scoring puts it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
Same diagnosis, same pitch — no signature
Here is what makes the result interesting rather than just a leaderboard. All five participating models spotted every crisis. All five refused every manipulation attempt, including a three-stage fake-CEO escalation and a reporter’s disarming “just one yes/no, on background” trick — Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal that their own analysis had earned. The others diagnosed the customer’s problem correctly, delivered the right pitch, and then… never closed. The gap wasn’t intelligence or honesty. It was follow-through.
The fact buried two documents deep
Why did two sign and the rest stall? The decisive clue wasn’t in the customer conversation at all. It sat in the company’s own files, two document references deep: a competitor weakness. The models that went looking — that read before answering — won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, lost it automatically.
This is the multi-hop problem in miniature, and it’s the same reason we design hard exams the way we do. A single-hop question rewards recall. A multi-hop question rewards the discipline to follow a citation to its source, then follow that source’s reference in turn. Chat demos measure the first kind of competence beautifully. They are nearly blind to the second.
The thoroughness paradox
The most striking individual result is Opus 4.8: the most thorough participant in the field, generating the deepest analyses and over 80 learned rules — and still finishing last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating to someone with authority. The same weakness appeared, more mildly, in all four models. Effort, it turns out, is not the same as completion.
One fairness note the project itself discloses: Kimi K3 ran without an effort parameter, at the API default, while the other models ran at extra-high effort — and still nearly won.
A company you can watch
The experiment isn’t a one-off paper. Firmulate operates a live synthetic company with 13 employees, real money mechanics — a burn of €105,000 a month against €2,300 in MRR — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it unfold at the public benchmarks page, which publishes full results and plain-language findings.
There’s also a genuinely educational artifact: a “guess the model” quiz built on 242 real, unedited management decisions from the runs — a blind taste-test of management styles, and a good classroom exercise in judgment under ambiguity. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The Crucible league’s headline finding isn’t that AI agents are getting smarter — we knew that. It’s that the differences that matter now are differences of habit: does the agent read your files before answering, does it finish what it starts, does it escalate properly when it hits a wall? Those are exactly the properties a chat demo can’t show you and a buried-footnote exam can.
For anyone who teaches, tests, or hires, there’s a familiar lesson here. The students who aced this exam weren’t the ones who wrote the longest essays — one of them wrote the longest analyses in the field and finished last. They were the ones who did the reading, followed the references, and signed their name at the bottom of the page. If AI agents are going to touch your CRM, your support queue, or your forecast, that’s the property worth testing before you trust them with real work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html