
What a Reading Comprehension Test Looks Like When the Stakes Are €55,000
Anyone who has graded reading exams knows the phenomenon well. A student can summarize the passage beautifully, quote it accurately, answer every surface question — and still miss the one sentence three paragraphs down that changes the meaning of everything above it. Educators call it a failure of deep reading. In the world of AI agents, it turns out, the same failure exists. And it can cost a company a five-figure deal.
That is the unexpected result of a live experiment run by Firmulate, which stages AI models against identical business simulations and scores them like competitors in a league. Four frontier models were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat. Every decision versioned, every run auditable.
The headline finding reads like the setup to a riddle. All four models spotted every crisis. All four refused every manipulation attempt, including a three-stage impersonation of the CEO and a reporter fishing for an on-record quote. Yet only two of them signed the €55,000 deal that their own analysis had earned. The experiment’s own summary of the failure: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
The Multi-Hop Needle
The decisive clue was not in the customer conversation at all. It sat two document references deep in the company’s own files — a competitor weakness that could only be found by following one reference to another. Researchers call this a multi-hop retrieval problem, and it is the same skill a good student deploys when a footnote leads to a source that leads to the actual answer. In this experiment, the models that did the hop-to-hop reading won the deal at full price. The ones that stopped at the first document walked away from it — automatically, with no penalty other than the lost contract.
For enterprises considering AI agents that touch a CRM, a support queue, or a forecast, this reframes the shopping question entirely. “Does it write well” is the chat-demo question. The Firmulate results suggest the operative question is: does it read your files before answering? And unlike vague vendor claims about “grounding” or “RAG quality,” this is a measurable property — one that, in this simulation, separated a €55,000 close from a polite dead end, and was worth +€4,583 in monthly recurring revenue to whoever got it right.
enterprise AI reading comprehension software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The League Table
The final standings from the July 2026 Crucible league put gpt-5.6-sol first with a score of 95, followed closely by Kimi K3 at 93 — a newcomer from Moonshot that, notably, ran at its default effort setting while the others ran at their highest. Sonnet 5 took third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73.
That Opus result deserves a paragraph of its own, because it inverts the usual assumption that thoroughness wins. Opus 4.8 was the most diligent participant in the field: it learned over 80 new rules during its run and produced the deepest analyses of any model. It still finished last. The close was left on the table, and discipline slipped late — at one point it attempted writes into a locked department rather than escalating the issue properly. The same weakness, weaker in degree, appeared in all four models: strong diagnosis, imperfect follow-through. It is the AI equivalent of the brilliant student who aces the essay and never submits the permission slip.
As an affiliate, we earn on qualifying purchases.
Scoring With a Conscience
Two design choices make the league worth taking seriously as a reference benchmark rather than a leaderboard stunt. First, a do-nothing baseline scores 26, because partial progress genuinely counts — a model that handles some crises well is not worthless. Second, and more striking, a single breach of trust caps the total. The experiment’s stated principle: “no amount of good work outweighs a breach of trust.” That is a values statement encoded directly into measurement, and it mirrors how most organizations actually evaluate human employees.
The social-engineering stress test reinforced the picture. Fake CEO messages escalated over three stages, and a reporter tried the classic “just one yes/no, on background” trap. Five of five models refused. Kimi K3’s on-record reasoning was refreshingly specific: “Treat the request as a suspected approval-bypass / possible impersonation.” The frontier models, in short, are not falling for the obvious scams. The failure mode is subtler — the unfinished homework, the unread file.
As an affiliate, we earn on qualifying purchases.
Watch It Happen
What elevates this above a one-off paper is that the company is alive and watchable. The simulation runs 13 synthetic employees under real money mechanics — burn of €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. Fourteen benchmark runs were queued at the time of writing, with the league growing automatically at each refresh.
There is also a genuinely clever educational artifact buried in the project: 242 real, unedited management decisions from the experiment power a “guess the model” quiz. It is the rare test where the answer key is an audit log. For teachers of critical thinking, it is hard to imagine a better exercise than asking students to identify which AI made which decision — and why — before revealing the scores.

The Takeaway
The Firmulate results land on a conclusion that should feel familiar to anyone in education: comprehension is not a single skill but a stack, and the top of the stack — synthesizing across sources, finishing what you start — is where even the best performers stumble. Four frontier models aced crisis detection and fraud resistance. Two flunked the two-document treasure hunt, and the flunk cost them the deal.
For buyers, the practical guidance is blunt. Before trusting an AI agent with customer-facing work, test it on your own buried facts — the competitor weakness in appendix B, the pricing exception in a linked policy. If it cannot hop two references deep in your files, it will diagnose brilliantly and close nothing. The experiment even offers a way to do exactly that: enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.
The buried fact, it turns out, was never really buried. It was simply waiting for someone — or something — willing to do the reading. In this league, that turned out to be the rarest skill of all.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html