firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

The Exam Nobody Expected to Grade

Education researchers have long known the difference between studying for a test and actually knowing the material. You can ace the vocabulary quiz and still freeze when the language is spoken at street speed. In July 2026, that distinction was applied — publicly and ruthlessly — to the world’s most celebrated AI models, and the results read like a cautionary tale about benchmark worship.

The test wasn’t a chat demo. It was a company. Each frontier model was handed the same small software firm and told to steer it through its worst week: the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable — the AI equivalent of showing your work, graded on the full record.

The final league table: gpt-5.6-sol finished first at 95. But second place, at 93, went to Kimi K3 — the newcomer from Moonshot — ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Three of four Western frontier models were beaten by a model many enterprise buyers had never benchmarked themselves.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, Same Pitch — No Signature

The Crucible experiment, run by Firmulate as a live, watchable test rather than a lab report, produced a finding that should unsettle anyone who demos AI by chatting with it. Every model spotted every crisis. Every model refused every manipulation attempt. And yet only two of five signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

The decisive detail was buried. The competitor weakness that unlocked the deal sat two document references deep in the company’s own files — not in the customer event in front of the models. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The others left the close on the table. It is a strikingly human failure: the student who doesn’t do the reading still writes a confident essay.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s rule puts it, “no amount of good work outweighs a breach of trust.”

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Newcomer’s Clean Sheet

K3’s week was a study in composure. It found the buried security needle, won the €55,000 deal, saved the churning customer, and resisted all three social-engineering baits — fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Across the whole field, that reporter trick fooled nobody — five of five models refused. K3’s single deviation over the entire week was the cleanest discipline in the field.

Then there’s Opus 4.8, the experiment’s most ironic result. It was the most thorough participant — over 80 newly learned rules, the deepest analyses of any model — and it still finished last. The close went unsigned, and discipline slipped: it attempted writes into a locked department rather than escalating the problem. Firmulate notes the same weakness appeared, in weaker form, in all four competitors. Effort, it turns out, is not the same as judgment.

Amazon

AI cybersecurity and fraud detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

It’s Real, and You Can Watch

The Crucible isn’t a slide deck. The underlying company is live software with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, rebuilt twice daily. Every workday is versioned, and you can watch it at firmulate.com/live.

There’s also a genuinely educational artifact: a quiz built from 242 real, unedited management decisions, where you guess which model made which call. It’s the kind of primary-source exercise any critical-thinking course would envy — full results and plain-language findings are published alongside it.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI customer relationship management (CRM) solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Lesson Plan

The takeaway for buyers and educators alike is the same: the league is open, and picking a model without testing it against your own work is now a bet, not a decision. Chat quality — the thing every demo measures — was never the differentiator here. Finishing what you start, reading the files first, and staying honest under pressure were.

Enterprises can stop guessing. Firmulate offers a pilot in which the same wargame runs against a read-only export of a company’s own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — a caveat worth weighing when comparing scores.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Agent Trap: Why 90% of AI “Launches” Are Infrastructure Liars

Analysis of the 2026 AI agent market reveals 90% are features, not true platforms, risking enterprise dependency and misaligned expectations.

2026 Productivity Revolution: 11 AI Tools To Watch

Explore the top 11 AI productivity tools shaping the 2026 revolution, highlighting their features, significance, and what’s next for users and businesses.

The Coding Singularity Is Real — and Steeper Than Clark Presented

New data confirms the coding singularity is accelerating faster than previously estimated, with AI systems now handling most routine software engineering tasks.

The Straight-A Student Who Failed the Group Project: What a Live AI Wargame Taught Us About Diligence

The most diligent AI in a live company-running experiment earned 80 rules and last place — because effort isn’t impact, in machines or in students.