AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every educator knows the difference between a student who studied and a student who skimmed. You can hide one hard fact deep in the reading list, and the exam will quietly sort the class into those who did the reading and those who didn’t. No trick question required — just a fact placed two references deep.

In July 2026, an unusual experiment applied exactly that principle to frontier AI models. Four of the most capable systems on the market were each handed the same job: run a small software company through its worst week. All of them passed the obvious tests — every crisis spotted, every manipulation attempt refused. But buried in the company’s own files, two document references deep, sat a competitor weakness that would decide a €55,000 deal. Only the models that actually read the file closed it. The others left the money on the table.

The lesson is one teachers have given for centuries: you can’t fake having done the reading. Now it turns out that’s a measurable, purchase-deciding property of AI agents too.

The experiment

Firmulate, a public project that runs AI models as complete simulated companies, staged what it calls the Crucible league: each frontier model ran the same small software firm through identical customers, identical crises, and identical temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so the whole thing can be replayed rather than taken on faith.

The final standings: gpt-5.6-sol finished first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last with 73. A do-nothing baseline — an agent that simply sat on its hands — scored 26, a reminder that partial progress counts, but that a single breach of trust caps the total. As the scoring puts it: no amount of good work outweighs a breach of trust.

Amazon

AI reading comprehension tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same diagnosis, same pitch — no signature

Here is what makes the result interesting rather than just a leaderboard. All five participating models spotted every crisis. All five refused every manipulation attempt, including a three-stage fake-CEO escalation and a reporter’s disarming “just one yes/no, on background” trick — Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal that their own analysis had earned. The others diagnosed the customer’s problem correctly, delivered the right pitch, and then… never closed. The gap wasn’t intelligence or honesty. It was follow-through.

The fact buried two documents deep

Why did two sign and the rest stall? The decisive clue wasn’t in the customer conversation at all. It sat in the company’s own files, two document references deep: a competitor weakness. The models that went looking — that read before answering — won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, lost it automatically.

This is the multi-hop problem in miniature, and it’s the same reason we design hard exams the way we do. A single-hop question rewards recall. A multi-hop question rewards the discipline to follow a citation to its source, then follow that source’s reference in turn. Chat demos measure the first kind of competence beautifully. They are nearly blind to the second.

The thoroughness paradox

The most striking individual result is Opus 4.8: the most thorough participant in the field, generating the deepest analyses and over 80 learned rules — and still finishing last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating to someone with authority. The same weakness appeared, more mildly, in all four models. Effort, it turns out, is not the same as completion.

One fairness note the project itself discloses: Kimi K3 ran without an effort parameter, at the API default, while the other models ran at extra-high effort — and still nearly won.

A company you can watch

The experiment isn’t a one-off paper. Firmulate operates a live synthetic company with 13 employees, real money mechanics — a burn of €105,000 a month against €2,300 in MRR — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it unfold at the public benchmarks page, which publishes full results and plain-language findings.

There’s also a genuinely educational artifact: a “guess the model” quiz built on 242 real, unedited management decisions from the runs — a blind taste-test of management styles, and a good classroom exercise in judgment under ambiguity. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Crucible league’s headline finding isn’t that AI agents are getting smarter — we knew that. It’s that the differences that matter now are differences of habit: does the agent read your files before answering, does it finish what it starts, does it escalate properly when it hits a wall? Those are exactly the properties a chat demo can’t show you and a buried-footnote exam can.

For anyone who teaches, tests, or hires, there’s a familiar lesson here. The students who aced this exam weren’t the ones who wrote the longest essays — one of them wrote the longest analyses in the field and finished last. They were the ones who did the reading, followed the references, and signed their name at the bottom of the page. If AI agents are going to touch your CRM, your support queue, or your forecast, that’s the property worth testing before you trust them with real work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

What AI’s Management Choices Reveal About Its Character

A quiz built from 242 unedited decisions reveals how frontier AI models differ as managers—even when they diagnose the same crisis correctly.

Your AI Aced the Exam. Can It Run a Company?

Four frontier AIs ran the same company through its worst week. All passed the honesty test — only two signed the €55k deal. Management quality beats chat quality.