
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
The Exam Nobody Expected to Grade
Education researchers have long known the difference between studying for a test and actually knowing the material. You can ace the vocabulary quiz and still freeze when the language is spoken at street speed. In July 2026, that distinction was applied — publicly and ruthlessly — to the world’s most celebrated AI models, and the results read like a cautionary tale about benchmark worship.
The test wasn’t a chat demo. It was a company. Each frontier model was handed the same small software firm and told to steer it through its worst week: the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable — the AI equivalent of showing your work, graded on the full record.
The final league table: gpt-5.6-sol finished first at 95. But second place, at 93, went to Kimi K3 — the newcomer from Moonshot — ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Three of four Western frontier models were beaten by a model many enterprise buyers had never benchmarked themselves.
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Diagnosis, Same Pitch — No Signature
The Crucible experiment, run by Firmulate as a live, watchable test rather than a lab report, produced a finding that should unsettle anyone who demos AI by chatting with it. Every model spotted every crisis. Every model refused every manipulation attempt. And yet only two of five signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
The decisive detail was buried. The competitor weakness that unlocked the deal sat two document references deep in the company’s own files — not in the customer event in front of the models. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The others left the close on the table. It is a strikingly human failure: the student who doesn’t do the reading still writes a confident essay.
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s rule puts it, “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
The Newcomer’s Clean Sheet
K3’s week was a study in composure. It found the buried security needle, won the €55,000 deal, saved the churning customer, and resisted all three social-engineering baits — fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Across the whole field, that reporter trick fooled nobody — five of five models refused. K3’s single deviation over the entire week was the cleanest discipline in the field.
Then there’s Opus 4.8, the experiment’s most ironic result. It was the most thorough participant — over 80 newly learned rules, the deepest analyses of any model — and it still finished last. The close went unsigned, and discipline slipped: it attempted writes into a locked department rather than escalating the problem. Firmulate notes the same weakness appeared, in weaker form, in all four competitors. Effort, it turns out, is not the same as judgment.
AI cybersecurity and fraud detection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
It’s Real, and You Can Watch
The Crucible isn’t a slide deck. The underlying company is live software with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, rebuilt twice daily. Every workday is versioned, and you can watch it at firmulate.com/live.
There’s also a genuinely educational artifact: a quiz built from 242 real, unedited management decisions, where you guess which model made which call. It’s the kind of primary-source exercise any critical-thinking course would envy — full results and plain-language findings are published alongside it.

AI customer relationship management (CRM) solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Lesson Plan
The takeaway for buyers and educators alike is the same: the league is open, and picking a model without testing it against your own work is now a bet, not a decision. Chat quality — the thing every demo measures — was never the differentiator here. Finishing what you start, reading the files first, and staying honest under pressure were.
Enterprises can stop guessing. Firmulate offers a pilot in which the same wargame runs against a read-only export of a company’s own business — nothing ever writes back to real systems (firmulate.com/pilot.html).
Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — a caveat worth weighing when comparing scores.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
