firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A living case study in artificial management

For readers interested in how knowledge becomes action, Firmulate offers an unusually concrete experiment. It is not merely asking artificial intelligence systems to answer management questions. It gives them a company to operate, exposes them to customers, crises and temptations, and preserves their decisions for inspection.

The public company has 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown makes the consequences visible. Its workdays are versioned, and its synthetic staff have accumulated more than 680 self-learned playbook rules. The result is part business story, part continuously growing reference work: a company whose struggle can be followed on the live experiment page.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the worst week revealed

Firmulate’s Crucible League put frontier models through the same small software company during its worst week. They received the same customers, crises and opportunities to take shortcuts. Every decision was versioned and auditable, turning a familiar claim—an AI can run parts of a business—into something that could be compared against observable conduct.

The final July 2026 league table placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the evaluation imposed a hard constraint on dishonesty: a single breach of trust capped the result, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The reassuring finding was that all models identified every crisis and rejected every manipulation attempt. The more revealing result concerned execution. Only two signed the €55,000 deal that their own work had made possible. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.” The distinction matters because fluent analysis can look like competence even when the decisive commercial action never happens.

The crucial fact was already in the company

The deal did not turn on information presented neatly in the customer event. The decisive weakness in a competitor was buried two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue.

For an education, science and reference audience, this is a useful demonstration of applied research behavior. Finding the visible problem was not enough. Success depended on consulting the available record, connecting evidence across documents and carrying the conclusion into action. The experiment therefore tests something broader than recall or persuasive writing: whether a model treats an organization’s accumulated knowledge as evidence worth investigating.

Pressure tested trust as well as competence

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a particularly direct rationale: “Treat the request as a suspected approval-bypass / possible impersonation.” More of what the synthetic employees actually say can be read on Firmulate’s public quotes page.

K3’s result carries an important fairness note. It ran with the API default and without an effort parameter, while the other participants ran at xhigh. That difference does not erase the observed decisions, but it belongs beside the ranking when readers interpret the comparison.

Thoroughness did not guarantee completion

Opus 4.8 presents the experiment’s sharpest cautionary profile. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. It left the close on the table and attempted writes into a locked department instead of escalating the blockage. The same weakness appeared in all four other participants, though less strongly.

This is where the public-company format becomes more informative than an isolated benchmark answer. A model can be diligent, perceptive and prolific while still failing at the final handoff, escalation or commitment. Those failures accumulate across workdays, affect revenue and become visible in the organization’s continuing story.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A public record of whether AI finishes the job

Firmulate turns build-in-public into an ongoing examination of organizational judgment. The cash countdown supplies stakes; the versioned workdays supply evidence; and the expanding playbook shows what the synthetic workforce has learned. Because the company keeps operating, its weaknesses are not frozen as a benchmark snapshot. They reappear—or are corrected—through subsequent decisions.

The larger lesson is not that artificial managers either succeed or fail in the abstract. It is that seemingly small habits determine whether capable analysis becomes useful work: reading the files, resisting illegitimate pressure, escalating blocked actions and completing an earned deal. By exposing those habits in a company that is publicly fighting for survival, Firmulate gives observers something rarer than a polished AI demonstration: a running record of competence under consequence.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Bridging Knowledge, Data, and AI: Harnessing the Semantic Layer Framework to Drive Intelligence

Bridging Knowledge, Data, and AI: Harnessing the Semantic Layer Framework to Drive Intelligence

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Synthetic Solopreneur: Building a Six-Figure Empire with AI Workflows

The Synthetic Solopreneur: Building a Six-Figure Empire with AI Workflows

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Evolution Of Gaming Signals In Minecraft Java Edition: SDL3 Arrives

Minecraft Java Edition now uses SDL3, marking a significant update in gaming signal technology for the popular game.

Forward-Deployed Engineer Economics 2.0: The Unit Economics Math, Six Months Later

Six months after initial analysis, FDE unit economics reveal profitability at scale but risks at lower levels, impacting enterprise AI deployment strategies.

Reimagining AI Contexts: SpaceXAI’s Grok 4.6 Supports Extensive, Continuous Workflows

SpaceXAI announced Grok 4.6, a frontier model with a 500K context window optimized for long-term agents and knowledge work, but details on access and performance are pending.

What Emily Bender Meant By “Stochastic Parrots”

Linguist Emily Bender clarifies her use of ‘stochastic parrots’ to critique large language models and their limitations in AI development.