
What can a management quiz teach us about artificial intelligence?
Most comparisons of frontier AI begin with knowledge: which model answers correctly, reasons most clearly or writes the strongest response? Firmulate asks a more practical question. What happens when the models must manage a company, face pressure and turn analysis into action?
The result is an unusually revealing educational experiment—and now an interactive one. A guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how an AI handled a situation and try to identify it from the decision itself. The challenge is entertaining, but it also teaches a serious lesson: models that appear similarly capable can develop recognizably different management personalities.
AI management decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, crises and temptations
Each frontier model was asked to run the same small software company through its worst week. The customers did not change. Neither did the crises or the temptations. Every decision was versioned and auditable, making the comparison less like a collection of polished demonstrations and more like a controlled management wargame.
The company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the consequences visible. Its models have accumulated more than 680 self-learned playbook rules, and every workday is versioned.
The striking result was not that one model noticed a crisis another missed. All the models spotted every crisis and refused every manipulation attempt. Their differences emerged in what happened afterward. Only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap crisply: “Same diagnosis, same pitch — no signature.”
The fact hidden in plain sight
The decisive competitive weakness was not contained in the customer event that triggered the work. It was buried two document references deep in the company’s own files. Models that followed the trail found it, won the deal at full price and added €4,583 in monthly recurring revenue.
That episode turns file-reading into a management competency rather than a clerical one. A model can interpret the immediate situation correctly and still miss the information that changes the commercial outcome. In this experiment, thorough use of the company’s existing knowledge separated diagnosis from execution.
Trust held under pressure
The models also faced fake messages from the chief executive that escalated across three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because the league’s scoring treats trust as a boundary, not a bonus. A do-nothing baseline scores 26 because partial progress counts, but a single breach of trust caps the total. The governing principle is explicit: “no amount of good work outweighs a breach of trust.”
A close finish with distinct failure modes
The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.
The ranking contains an instructive paradox. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four models.
There is also an important fairness qualification. Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That does not erase its result, but it belongs beside any interpretation of the narrow gap at the top.

AI file reading and knowledge management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From model intelligence to organizational judgment
For educators and technically curious readers, the quiz offers something more useful than AI trivia. It makes model behavior observable through concrete decisions: whether a manager reads the available evidence, protects trust, respects boundaries and completes commercially valuable work.
Firmulate’s larger proposition is that organizations can examine those traits before deploying an AI workforce. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to their real systems.
The live experiment shows why that evaluation cannot stop at fluent answers. Every participant recognized danger and resisted manipulation, yet their final performances ranged from 73 to 95. The decisive distinction was not merely knowing what should happen. It was reliably making it happen.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trust and decision-making evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI management personality assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.