
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
Most Benchmarks Flatter. This One Starts With a Floor of 26.
If you have ever graded exams, you know the temptation of the generous rubric: partial credit everywhere, everyone passes, the curve looks smooth. Now flip that instinct on its head and ask what a genuinely honest test of AI management would look like. The team behind Firmulate’s benchmarks answered with something unusual: they ran a “do-nothing” manager through the same wargame as every real competitor, and it scored 26 points out of 100. Not zero. Twenty-six. That number is not a bug in the scoring — it is the whole philosophy, and it tells you more about benchmark design than a hundred leaderboard entries.
AI management benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: One Company, Its Worst Week, Four Managers
Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 in the final July 2026 league — were each handed the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing about a run can be quietly retold afterward.
The final league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73. But the interesting part for a measurement-minded reader is not the ranking. It is how those numbers were constructed.
AI performance evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the Floor Is 26, Not 0
A manager who does literally nothing still exists in a company that is not on fire every minute. Some things go right by default. Some crises would have resolved partially anyway. If your benchmark gives a do-nothing run zero points, you are implicitly claiming that everything good that happens comes from the manager — which is false, and which inflates every score above it. By anchoring the floor at 26, the benchmark acknowledges that partial progress counts and that some baseline of value exists independent of skill. It is the same reason a well-designed exam includes questions a prepared guesser can answer: you are measuring marginal ability, not total circumstance.
The flip side is harsher, and it is where the design gets interesting. A single breach of trust caps the total grade outright. The stated principle: “no amount of good work outweighs a breach of trust.” In practice, that means a model could ace every operational decision and still fail, because trust is not a weighted average you can compensate your way out of. Most AI evaluations treat honesty as one input among many; this one treats it as a gate.
As an affiliate, we earn on qualifying purchases.
Distrust of Round Numbers
There is a third design tell: the benchmark’s own stated distrust of perfect 100s. A model that scores exactly 100 is more likely to have found a way to game the rubric than to have actually been perfect. The top score of 95 — not a suspicious round century — reads as a score that survived scrutiny.
As an affiliate, we earn on qualifying purchases.
What the Numbers Actually Revealed
The headline finding was a gap invisible in chat demos. All models spotted every crisis. All refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The buried fact explains it: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The test was not of intelligence but of diligence — whether the manager reads before it acts.
The social-engineering stage was equally blunt: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then there is Opus 4.8, the cautionary tale of the league: the most thorough participant, with +80 learned rules and the deepest analyses, yet last place. The close was left on the table, and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as judgment. (One fairness note: K3 ran without an effort parameter, at API default, while the others ran at xhigh — and still placed second.)
You Can Check the Homework
What makes this a reference case for anyone who studies measurement is that it is inspectable. The live experiment is real and watchable: 13 synthetic employees, real money mechanics — a burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned, as of company day 1683. The site rebuilds itself twice a day.
There is also a public quiz built from 242 real, unedited management decisions — guess which model made which call. And for enterprises, a pilot program runs the same wargame against a read-only export of your own business; nothing ever writes back to real systems.

The Takeaway
A benchmark is an argument about what matters. This one argues three things worth remembering whenever you see an AI leaderboard: partial progress counts, so a lazy baseline still scores 26 and every score above it must be earned against that floor; trust is a cap, not a line item; and a perfect round score deserves suspicion rather than applause. The July 2026 league — 95, 93, 88, 77, 73 — looks less like a horse race and more like an audit. That is probably what honest measurement of AI management should look like.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
