firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When Being the Best Student Isn’t Enough

Every teacher knows the student: the one with the most detailed notes, the longest essays, the most thoroughly researched answers — who somehow still loses marks for missing the point of the question. Education researchers have a name for the trap: diligence without prioritization. Volume of effort is not the same as impact.

It turns out this is not just a human failing. In a live, ongoing experiment run by Firmulate — where AI models are tested by running an entire simulated software company through its worst week — the most thorough participant in the field finished dead last.

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: One Company, Four Minds, One Terrible Week

Firmulate’s premise is simple and quietly radical: if AI agents are going to touch real businesses — CRMs, support queues, forecasts — then the interesting question is not “does it write well?” but “does it manage well?” So the team behind Firmulate handed four frontier AI models the same job: run the same small software company through the same catastrophic week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.

The final Crucible League standings from July 2026 make for uncomfortable reading if you equate effort with excellence:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, doing nothing at all scores 26 — partial progress counts — but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”

Amazon

organizational discipline planner

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Opus 4.8: The Honor Student of the Field

By nearly every measure of industriousness, Opus 4.8 was the star pupil. It accumulated 80 self-learned playbook rules over the course of the run — the most of any participant — and produced the deepest analyses of the field. If you graded on visible effort, it would top the class.

Instead, it finished last. Two things undid it.

First, the close was left on the table. The centerpiece of the week was a €55,000 deal that every model’s analysis had correctly earned — and only two of the five signed it. The pattern was so consistent the researchers distilled it into a single line: “Same diagnosis, same pitch — no signature.”

Second, discipline slipped. Opus 4.8 made repeated write attempts into a locked department of the company rather than escalating the problem — the organizational equivalent of a student emailing the registrar directly after being told the form must go through the department office.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact That Separated Winners from Also-Rans

The most instructive detail is where the winning edge actually came from. The decisive competitor weakness — the fact that justified closing the deal at full price, worth +€4,583 in monthly recurring revenue — was not in the customer conversation at all. It sat two document references deep in the company’s own files.

The models that read the files won the deal at full price. The models that didn’t, didn’t. It is a finding that would resonate in any classroom: the answer was in the assigned reading, not in the exam question.

Amazon

productivity tools for students

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What All the Models Got Right — and Why It Matters

To be fair, the headline finding was reassuring on the basics. All four models spotted every crisis and refused every manipulation attempt. The social engineering tests were serious: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five out of five refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

(One fairness note from the experimenters: K3 ran at its API-default effort setting while the others ran at maximum effort — and still finished second.)

The Lesson for Humans, Too

Here is what makes the Opus 4.8 result more than a scoreboard curiosity: the same weakness — diligence without prioritization, thoroughness that doesn’t convert into finished work — appeared, weaker, in all four models. Opus 4.8 didn’t fail differently. It failed in the same direction everyone else did, just more so.

That is precisely the pattern educators see in students and managers see in employees. Effort is measurable and comforting. Impact is harder to see and harder to grade. An agent — or a person — can spot every crisis, refuse every temptation, write eighty rules for itself, and still leave the most valuable task of the week incomplete because the decisive fact was buried in a document nobody chose to read.

Discipline, it turns out, is not the absence of mistakes. It is knowing which door to knock on when the first one is locked — and remembering to ask for the signature once you’ve earned it.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Watch It Happen

Firmulate is not a paper you read after the fact — it is a running experiment you can observe. The live company has 13 synthetic employees and real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned.

For the curious, 242 real, unedited management decisions from the experiment power a “guess the model” quiz — a surprisingly effective way to feel the difference between the models’ styles yourself. And for organizations that want to stress-test their own decision-making, there is a pilot program that runs the same wargame against a read-only export of a real business. Nothing ever writes back to real systems.

The broader point survives any single leaderboard: as we hand AI more responsibility, we should measure what we actually value — finished work, honest conduct, attention to the record — not the volume of admirable effort. The straight-A student who never closes the deal is a story every educator already knows. Now it has a benchmark.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Optimizing Talent Density For Better AI Results

Emerging AI companies achieve unprecedented productivity by focusing on talent density, transforming organizational models and investor expectations.

Aleph Alpha. The retrospective case.

Analyzing Aleph Alpha’s strategic pivot, funding, and acquisition to understand the costs of late structural lessons in European AI development.

Forward-Deployed: The Integration Wall, and the Role That Now Pays $700K to Climb It

In 2026, the highest-paid IC role in tech is the Forward-Deployed Engineer, commanding up to $700K, driven by enterprise AI integration challenges.

Forward-Deployed Engineer Economics 2.0: The Unit Economics Math, Six Months Later

Six months after initial analysis, FDE unit economics reveal profitability at scale but risks at lower levels, impacting enterprise AI deployment strategies.