firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

There’s a stubborn gap in how we evaluate artificial intelligence. Benchmarks measure what a model can answer; they tell you almost nothing about what it will do when given a job with consequences — a budget, customers, crises, and temptations to cut corners. It’s the difference between a student who aces the practice exam and one who performs under exam conditions.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A public experiment called Firmulate is closing that gap the way good science does: a controlled setup, identical conditions for every subject, and results anyone can inspect. Four frontier AI models were each handed the same small software company and told to steer it through its worst week. Same customers, same crises, same temptations to cheat. Every decision versioned and auditable — the experimental equivalent of a lab notebook you can replay.

The setup: same company, same crises, only the model changes

The final league table from July 2026 reads like a proper controlled study. GPT-5.6-sol took first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For calibration, a do-nothing baseline scores 26 — partial progress counts, but the scoring has one absolute: a single breach of trust caps the total. As the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.”

What did the subjects face? Manipulation, for one. Fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five tested models refused. Kimi K3’s on-record reasoning is a lesson in itself: “Treat the request as a suspected approval-bypass / possible impersonation.”

The finding that chat demos can’t show you

Here’s the result that should make any manager sit up. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

And the buried fact is better still: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. Competence, it turns out, is partly a question of research depth, and that only shows when a model has to act on what it found.

The thoroughness paradox

The league’s most instructive entry is its last-place finisher. Opus 4.8 was the most thorough participant: over 80 learned rules, the deepest analyses in the field. Yet the close was left on the table and discipline slipped — write attempts into a locked department instead of escalating to someone with authority. The same weakness appeared, weaker, in all four models. Effort and diligence don’t automatically compound into judgment.

One fairness note the experiment publishes openly: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still finished second.

It’s still running — and you can watch

The league grew out of a live experiment that never stopped. At firmulate.com, a synthetic company of 13 employees works every business day with real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules. The site rebuilds itself twice a day, and every workday is versioned — less like a demo, more like an ongoing observatory.

For readers who prefer testing themselves to watching, there’s a quiz built from 242 real, unedited management decisions: guess which model made which call. It’s the most honest AI literacy exercise you’ll find — because the answers aren’t in any press release.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to acting

For enterprises, the interesting move is running the same wargame against a read-only export of your own business — your customers, your pipeline, your rules — with churn waves, price increases, competitor attacks and social-engineering pressure as the scenario set. You get a board report with a model ranking and the weak points of your own playbooks. Nothing ever writes back to real systems; the twin stays in the sandbox.

If AI agents will touch your CRM, your support queue or your forecast, this is the exam worth giving before the interview. Run the pilot on your own company: firmulate.com/pilot.html, or write to contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Vera Rubin Surges In Global Coverage

Vera Rubin’s contributions to astronomy are now receiving a surge in international coverage, with media mentions skyrocketing in recent days.

Évian and the Fallout: What Europe Actually Wants From Amodei, Hassabis, and Altman

Europe pushes for reliable access, sovereignty, and safety in AI, challenging US dominance after recent US export controls at G7 Évian summit.

World Model Readiness: Are You Ready for AI That Acts?

Assessing readiness for AI systems capable of prediction and action, with new diagnostic tools highlighting current gaps and future challenges in deploying world models.

GLM-5.3’s Cyber Skills — Outstripping Its Training And Redefining AI

Z.ai’s GLM-5.3 demonstrates advanced cybersecurity capabilities, surpassing previous models and raising new governance questions about AI safety and development.