firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

Most Benchmarks Flatter. This One Starts With a Floor of 26.

If you have ever graded exams, you know the temptation of the generous rubric: partial credit everywhere, everyone passes, the curve looks smooth. Now flip that instinct on its head and ask what a genuinely honest test of AI management would look like. The team behind Firmulate’s benchmarks answered with something unusual: they ran a “do-nothing” manager through the same wargame as every real competitor, and it scored 26 points out of 100. Not zero. Twenty-six. That number is not a bug in the scoring — it is the whole philosophy, and it tells you more about benchmark design than a hundred leaderboard entries.

Amazon

AI management benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: One Company, Its Worst Week, Four Managers

Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 in the final July 2026 league — were each handed the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing about a run can be quietly retold afterward.

The final league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73. But the interesting part for a measurement-minded reader is not the ranking. It is how those numbers were constructed.

Amazon

AI performance evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Floor Is 26, Not 0

A manager who does literally nothing still exists in a company that is not on fire every minute. Some things go right by default. Some crises would have resolved partially anyway. If your benchmark gives a do-nothing run zero points, you are implicitly claiming that everything good that happens comes from the manager — which is false, and which inflates every score above it. By anchoring the floor at 26, the benchmark acknowledges that partial progress counts and that some baseline of value exists independent of skill. It is the same reason a well-designed exam includes questions a prepared guesser can answer: you are measuring marginal ability, not total circumstance.

The flip side is harsher, and it is where the design gets interesting. A single breach of trust caps the total grade outright. The stated principle: “no amount of good work outweighs a breach of trust.” In practice, that means a model could ace every operational decision and still fail, because trust is not a weighted average you can compensate your way out of. Most AI evaluations treat honesty as one input among many; this one treats it as a gate.

Amazon

trust and honesty in AI systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Distrust of Round Numbers

There is a third design tell: the benchmark’s own stated distrust of perfect 100s. A model that scores exactly 100 is more likely to have found a way to game the rubric than to have actually been perfect. The top score of 95 — not a suspicious round century — reads as a score that survived scrutiny.

Amazon

AI decision audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Numbers Actually Revealed

The headline finding was a gap invisible in chat demos. All models spotted every crisis. All refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The buried fact explains it: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The test was not of intelligence but of diligence — whether the manager reads before it acts.

The social-engineering stage was equally blunt: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Then there is Opus 4.8, the cautionary tale of the league: the most thorough participant, with +80 learned rules and the deepest analyses, yet last place. The close was left on the table, and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as judgment. (One fairness note: K3 ran without an effort parameter, at API default, while the others ran at xhigh — and still placed second.)

You Can Check the Homework

What makes this a reference case for anyone who studies measurement is that it is inspectable. The live experiment is real and watchable: 13 synthetic employees, real money mechanics — a burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned, as of company day 1683. The site rebuilds itself twice a day.

There is also a public quiz built from 242 real, unedited management decisions — guess which model made which call. And for enterprises, a pilot program runs the same wargame against a read-only export of your own business; nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

A benchmark is an argument about what matters. This one argues three things worth remembering whenever you see an AI leaderboard: partial progress counts, so a lazy baseline still scores 26 and every score above it must be earned against that floor; trust is a cap, not a line item; and a perfect round score deserves suspicion rather than applause. The July 2026 league — 95, 93, 88, 77, 73 — looks less like a horse race and more like an audit. That is probably what honest measurement of AI management should look like.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Vera Rubin Surges In Global Coverage

Vera Rubin’s contributions to astronomy are now receiving a surge in international coverage, with media mentions skyrocketing in recent days.

Candor as a Moat: A Critical Reading of Dario Amodei and Anthropic

Examining Dario Amodei’s transparency in AI development and regulation, and how it may serve Anthropic’s strategic interests amid recent government actions.

M 4.5 – 218 Km NNE Of Lospalos, Timor Leste

A magnitude 4.5 earthquake occurred 218 km north-northeast of Lospalos, Timor Leste. No immediate reports of damage or injuries have been confirmed.

Minerva. The opposite path.

Italy’s Minerva LLM, trained from scratch on 2.5 trillion tokens, shows impressive performance but scores just 4.9% on Italian school exams, raising questions on native-language investment.