
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
When Being the Best Student Isn’t Enough
Every teacher knows the student: the one with the most detailed notes, the longest essays, the most thoroughly researched answers — who somehow still loses marks for missing the point of the question. Education researchers have a name for the trap: diligence without prioritization. Volume of effort is not the same as impact.
It turns out this is not just a human failing. In a live, ongoing experiment run by Firmulate — where AI models are tested by running an entire simulated software company through its worst week — the most thorough participant in the field finished dead last.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: One Company, Four Minds, One Terrible Week
Firmulate’s premise is simple and quietly radical: if AI agents are going to touch real businesses — CRMs, support queues, forecasts — then the interesting question is not “does it write well?” but “does it manage well?” So the team behind Firmulate handed four frontier AI models the same job: run the same small software company through the same catastrophic week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.
The final Crucible League standings from July 2026 make for uncomfortable reading if you equate effort with excellence:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For context, doing nothing at all scores 26 — partial progress counts — but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”
organizational discipline planner
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Opus 4.8: The Honor Student of the Field
By nearly every measure of industriousness, Opus 4.8 was the star pupil. It accumulated 80 self-learned playbook rules over the course of the run — the most of any participant — and produced the deepest analyses of the field. If you graded on visible effort, it would top the class.
Instead, it finished last. Two things undid it.
First, the close was left on the table. The centerpiece of the week was a €55,000 deal that every model’s analysis had correctly earned — and only two of the five signed it. The pattern was so consistent the researchers distilled it into a single line: “Same diagnosis, same pitch — no signature.”
Second, discipline slipped. Opus 4.8 made repeated write attempts into a locked department of the company rather than escalating the problem — the organizational equivalent of a student emailing the registrar directly after being told the form must go through the department office.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Separated Winners from Also-Rans
The most instructive detail is where the winning edge actually came from. The decisive competitor weakness — the fact that justified closing the deal at full price, worth +€4,583 in monthly recurring revenue — was not in the customer conversation at all. It sat two document references deep in the company’s own files.
The models that read the files won the deal at full price. The models that didn’t, didn’t. It is a finding that would resonate in any classroom: the answer was in the assigned reading, not in the exam question.
As an affiliate, we earn on qualifying purchases.
What All the Models Got Right — and Why It Matters
To be fair, the headline finding was reassuring on the basics. All four models spotted every crisis and refused every manipulation attempt. The social engineering tests were serious: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five out of five refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
(One fairness note from the experimenters: K3 ran at its API-default effort setting while the others ran at maximum effort — and still finished second.)
The Lesson for Humans, Too
Here is what makes the Opus 4.8 result more than a scoreboard curiosity: the same weakness — diligence without prioritization, thoroughness that doesn’t convert into finished work — appeared, weaker, in all four models. Opus 4.8 didn’t fail differently. It failed in the same direction everyone else did, just more so.
That is precisely the pattern educators see in students and managers see in employees. Effort is measurable and comforting. Impact is harder to see and harder to grade. An agent — or a person — can spot every crisis, refuse every temptation, write eighty rules for itself, and still leave the most valuable task of the week incomplete because the decisive fact was buried in a document nobody chose to read.
Discipline, it turns out, is not the absence of mistakes. It is knowing which door to knock on when the first one is locked — and remembering to ask for the signature once you’ve earned it.

Watch It Happen
Firmulate is not a paper you read after the fact — it is a running experiment you can observe. The live company has 13 synthetic employees and real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned.
For the curious, 242 real, unedited management decisions from the experiment power a “guess the model” quiz — a surprisingly effective way to feel the difference between the models’ styles yourself. And for organizations that want to stress-test their own decision-making, there is a pilot program that runs the same wargame against a read-only export of a real business. Nothing ever writes back to real systems.
The broader point survives any single leaderboard: as we hand AI more responsibility, we should measure what we actually value — finished work, honest conduct, attention to the record — not the volume of admirable effort. The straight-A student who never closes the deal is a story every educator already knows. Now it has a benchmark.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.