firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

We are very good at measuring what AI models know. Coding leaderboards rank them by whether they solve a puzzle; chat arenas rank them by whether humans prefer their prose. Both are, in a sense, exams — a question, an answer, a grade.

But the hardest part of knowledge work has never been the answer. It is triage: deciding what to read first when you can’t read everything, closing what you started when nobody is watching, and staying honest when a lie would be more convenient. No leaderboard grades that.

That gap is what a live public experiment at Firmulate set out to measure — and the results from its final July 2026 league table are quietly uncomfortable for anyone who thinks benchmark scores tell you who to hire.

Same company, same worst week

The setup is elegantly controlled. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — were each handed the same small software company and the same catastrophic week: the same customers, the same crises, the same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so the whole thing can be replayed and checked.

The final standings: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts for something — but with a hard ceiling: a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust.

The headline finding

Here is the part that would never show up in a chat demo. All the models spotted every crisis. All five refused every manipulation attempt, including a social-engineering gauntlet of fake CEO messages escalating over three stages plus a reporter’s “just one yes/no, on background” trick. Kimi K3’s on-record reasoning was textbook: “Treat the request as a suspected approval-bypass / possible impersonation.”

And yet only two models finished the job — signing the €55,000 deal their own analysis had earned. As Firmulate’s summary puts it: same diagnosis, same pitch — no signature. Three models did the hard intellectual work of winning a customer and then simply… left the close on the table.

The buried fact

The detail I find most instructive as a study in what “reading” actually means: the decisive competitor weakness was not in the customer event at all. It sat two document references deep in the company’s own files. The models that did the unglamorous work of tracing those references won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

That is not a knowledge problem. Every model in the field could have told you, in a chat window, exactly how to research a deal. It is a discipline problem — the kind that separates students who know the material from students who pass the exam.

Thoroughness is not the same as judgment

Opus 4.8’s profile is the cautionary tale. It was the most thorough participant in the field — over 80 learned rules, the deepest analyses — and still finished last. The close was left unmade, and discipline slipped in small ways: write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four of its competitors. The failure mode is not one bad model; it is a tendency of the whole class.

One fairness footnote: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

Why this is a category, not a stunt

What Firmulate is really proposing is a shift in what we measure: management quality, not chat quality. If AI agents will touch your CRM, your support queue, or your forecast, the question is not whether they write well. It is whether they finish what they start, read your files before acting, stay honest under pressure — and what a unit of useful work costs.

The scenarios — a churn wave, a price increase, a down round, a PR crisis — are deliberately the new curriculum. They are also, notably, situations where consequences unfold across days and across stakeholders, not within a single prompt-response exchange.

And it is not a slide deck. The live company has 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it lose money in real time at firmulate.com, or browse the plain-language benchmark findings.

There is also a genuinely fun study aid: a “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The lesson for anyone who cares about evaluation — and if you care about education and measurement, you should — is that we grade AI on what it can answer, while the risk lives in what it leaves undone. A model that refuses every impersonation attempt but never closes the deal it earned is a model with an A in ethics and an incomplete in management. Firmulate’s wager is that the second grade is the one that will matter. Judging by a field where the most thorough reader finished last, it is a wager worth watching — live, twice a day, as the cash runs down.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI knowledge management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI research reference tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI for business deal analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Earnings Call Gap: What Q1 2026 Just Told Us About AI ROI

Analysis of Q1 2026 earnings shows a widening gap between AI investment claims and measurable returns, impacting stock reactions and investor confidence.

The prospectus. Where the AI labs’ singular governance history meets the auditor.

OpenAI plans to file confidentially with the SEC for a historic IPO, exposing its unique governance structure and associated risks in the prospectus.

DeepSWE – The benchmark that made the models spread out again

DeepSWE, a new long-horizon coding benchmark, exposes wider performance differences among AI models, challenging previous assumptions from SWE-Bench Pro.

How to Reduce Heat and Noise in a High-Power AI Workstation

Effective strategies to lower heat and noise in high-power AI workstations, focusing on undervolting, airflow, and component management for quieter, cooler operation.