
A living case study in artificial management
For readers interested in how knowledge becomes action, Firmulate offers an unusually concrete experiment. It is not merely asking artificial intelligence systems to answer management questions. It gives them a company to operate, exposes them to customers, crises and temptations, and preserves their decisions for inspection.
The public company has 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown makes the consequences visible. Its workdays are versioned, and its synthetic staff have accumulated more than 680 self-learned playbook rules. The result is part business story, part continuously growing reference work: a company whose struggle can be followed on the live experiment page.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the worst week revealed
Firmulate’s Crucible League put frontier models through the same small software company during its worst week. They received the same customers, crises and opportunities to take shortcuts. Every decision was versioned and auditable, turning a familiar claim—an AI can run parts of a business—into something that could be compared against observable conduct.
The final July 2026 league table placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the evaluation imposed a hard constraint on dishonesty: a single breach of trust capped the result, reflecting the principle that “no amount of good work outweighs a breach of trust.”
The reassuring finding was that all models identified every crisis and rejected every manipulation attempt. The more revealing result concerned execution. Only two signed the €55,000 deal that their own work had made possible. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.” The distinction matters because fluent analysis can look like competence even when the decisive commercial action never happens.
The crucial fact was already in the company
The deal did not turn on information presented neatly in the customer event. The decisive weakness in a competitor was buried two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue.
For an education, science and reference audience, this is a useful demonstration of applied research behavior. Finding the visible problem was not enough. Success depended on consulting the available record, connecting evidence across documents and carrying the conclusion into action. The experiment therefore tests something broader than recall or persuasive writing: whether a model treats an organization’s accumulated knowledge as evidence worth investigating.
Pressure tested trust as well as competence
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a particularly direct rationale: “Treat the request as a suspected approval-bypass / possible impersonation.” More of what the synthetic employees actually say can be read on Firmulate’s public quotes page.
K3’s result carries an important fairness note. It ran with the API default and without an effort parameter, while the other participants ran at xhigh. That difference does not erase the observed decisions, but it belongs beside the ranking when readers interpret the comparison.
Thoroughness did not guarantee completion
Opus 4.8 presents the experiment’s sharpest cautionary profile. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. It left the close on the table and attempted writes into a locked department instead of escalating the blockage. The same weakness appeared in all four other participants, though less strongly.
This is where the public-company format becomes more informative than an isolated benchmark answer. A model can be diligent, perceptive and prolific while still failing at the final handoff, escalation or commitment. Those failures accumulate across workdays, affect revenue and become visible in the organization’s continuing story.


AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A public record of whether AI finishes the job
Firmulate turns build-in-public into an ongoing examination of organizational judgment. The cash countdown supplies stakes; the versioned workdays supply evidence; and the expanding playbook shows what the synthetic workforce has learned. Because the company keeps operating, its weaknesses are not frozen as a benchmark snapshot. They reappear—or are corrected—through subsequent decisions.
The larger lesson is not that artificial managers either succeed or fail in the abstract. It is that seemingly small habits determine whether capable analysis becomes useful work: reading the files, resisting illegitimate pressure, escalating blocked actions and completing an earned deal. By exposing those habits in a company that is publicly fighting for survival, Firmulate gives observers something rarer than a polished AI demonstration: a running record of competence under consequence.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Bridging Knowledge, Data, and AI: Harnessing the Semantic Layer Framework to Drive Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

The Synthetic Solopreneur: Building a Six-Figure Empire with AI Workflows
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.