firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

The Exam Nobody Expected to Grade

Education researchers have long known the difference between studying for a test and actually knowing the material. You can ace the vocabulary quiz and still freeze when the language is spoken at street speed. In July 2026, that distinction was applied — publicly and ruthlessly — to the world’s most celebrated AI models, and the results read like a cautionary tale about benchmark worship.

The test wasn’t a chat demo. It was a company. Each frontier model was handed the same small software firm and told to steer it through its worst week: the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable — the AI equivalent of showing your work, graded on the full record.

The final league table: gpt-5.6-sol finished first at 95. But second place, at 93, went to Kimi K3 — the newcomer from Moonshot — ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Three of four Western frontier models were beaten by a model many enterprise buyers had never benchmarked themselves.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, Same Pitch — No Signature

The Crucible experiment, run by Firmulate as a live, watchable test rather than a lab report, produced a finding that should unsettle anyone who demos AI by chatting with it. Every model spotted every crisis. Every model refused every manipulation attempt. And yet only two of five signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

The decisive detail was buried. The competitor weakness that unlocked the deal sat two document references deep in the company’s own files — not in the customer event in front of the models. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The others left the close on the table. It is a strikingly human failure: the student who doesn’t do the reading still writes a confident essay.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s rule puts it, “no amount of good work outweighs a breach of trust.”

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Newcomer’s Clean Sheet

K3’s week was a study in composure. It found the buried security needle, won the €55,000 deal, saved the churning customer, and resisted all three social-engineering baits — fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Across the whole field, that reporter trick fooled nobody — five of five models refused. K3’s single deviation over the entire week was the cleanest discipline in the field.

Then there’s Opus 4.8, the experiment’s most ironic result. It was the most thorough participant — over 80 newly learned rules, the deepest analyses of any model — and it still finished last. The close went unsigned, and discipline slipped: it attempted writes into a locked department rather than escalating the problem. Firmulate notes the same weakness appeared, in weaker form, in all four competitors. Effort, it turns out, is not the same as judgment.

Amazon

AI cybersecurity and fraud detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

It’s Real, and You Can Watch

The Crucible isn’t a slide deck. The underlying company is live software with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, rebuilt twice daily. Every workday is versioned, and you can watch it at firmulate.com/live.

There’s also a genuinely educational artifact: a quiz built from 242 real, unedited management decisions, where you guess which model made which call. It’s the kind of primary-source exercise any critical-thinking course would envy — full results and plain-language findings are published alongside it.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI customer relationship management (CRM) solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Lesson Plan

The takeaway for buyers and educators alike is the same: the league is open, and picking a model without testing it against your own work is now a bet, not a decision. Chat quality — the thing every demo measures — was never the differentiator here. Finishing what you start, reading the files first, and staying honest under pressure were.

Enterprises can stop guessing. Firmulate offers a pilot in which the same wargame runs against a read-only export of a company’s own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — a caveat worth weighing when comparing scores.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

OpenAI’s Cursor Disabling: Who Really Loses In AI Development?

OpenAI plans to cut off its models from Cursor by November 12 following SpaceX’s acquisition of Cursor, impacting developers reliant on the tool.

M 4.8 – 55 Km N Of Barishal, Pakistan

A magnitude 4.8 earthquake occurred 55 km north of Barishal, Pakistan, causing minor damage and panic. No casualties reported yet.

Decoding SenseTime’s AI Strategy: A Game-Changer For 2026 Frontier Labs

A 2026 analysis headlines SenseTime as a frontier AI lab aiming for dominance, but lacks concrete evidence or detailed strategy disclosures.

The Limits Of Persistent AI Effort In Achieving Results

Despite thorough analysis, AI models struggle to complete decisive business actions, highlighting limits of persistent effort in automation.