firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

What can a management quiz teach us about artificial intelligence?

Most comparisons of frontier AI begin with knowledge: which model answers correctly, reasons most clearly or writes the strongest response? Firmulate asks a more practical question. What happens when the models must manage a company, face pressure and turn analysis into action?

The result is an unusually revealing educational experiment—and now an interactive one. A guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how an AI handled a situation and try to identify it from the decision itself. The challenge is entertaining, but it also teaches a serious lesson: models that appear similarly capable can develop recognizably different management personalities.

Amazon

AI management decision simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, crises and temptations

Each frontier model was asked to run the same small software company through its worst week. The customers did not change. Neither did the crises or the temptations. Every decision was versioned and auditable, making the comparison less like a collection of polished demonstrations and more like a controlled management wargame.

The company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the consequences visible. Its models have accumulated more than 680 self-learned playbook rules, and every workday is versioned.

The striking result was not that one model noticed a crisis another missed. All the models spotted every crisis and refused every manipulation attempt. Their differences emerged in what happened afterward. Only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap crisply: “Same diagnosis, same pitch — no signature.”

The fact hidden in plain sight

The decisive competitive weakness was not contained in the customer event that triggered the work. It was buried two document references deep in the company’s own files. Models that followed the trail found it, won the deal at full price and added €4,583 in monthly recurring revenue.

That episode turns file-reading into a management competency rather than a clerical one. A model can interpret the immediate situation correctly and still miss the information that changes the commercial outcome. In this experiment, thorough use of the company’s existing knowledge separated diagnosis from execution.

Trust held under pressure

The models also faced fake messages from the chief executive that escalated across three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because the league’s scoring treats trust as a boundary, not a bonus. A do-nothing baseline scores 26 because partial progress counts, but a single breach of trust caps the total. The governing principle is explicit: “no amount of good work outweighs a breach of trust.”

A close finish with distinct failure modes

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

The ranking contains an instructive paradox. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four models.

There is also an important fairness qualification. Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That does not erase its result, but it belongs beside any interpretation of the narrow gap at the top.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI file reading and knowledge management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From model intelligence to organizational judgment

For educators and technically curious readers, the quiz offers something more useful than AI trivia. It makes model behavior observable through concrete decisions: whether a manager reads the available evidence, protects trust, respects boundaries and completes commercially valuable work.

Firmulate’s larger proposition is that organizations can examine those traits before deploying an AI workforce. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to their real systems.

The live experiment shows why that evaluation cannot stop at fluent answers. Every participant recognized danger and resisted manipulation, yet their final performances ranged from 73 to 95. The decisive distinction was not merely knowing what should happen. It was reliably making it happen.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and decision-making evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI management personality assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Revolution: Microsoft’s Signal Peak 2026 And The Anthropic Alliance

Microsoft’s Signal Peak 2026, a new AI security platform, integrates Anthropic’s models, signaling a major shift in enterprise AI and security strategies.

Breakdown Of AI Memory: The 176GB You Never Read About

Exploring the overlooked memory costs in AI models, focusing on the 176GB weight size and the critical role of the KV cache in model performance.

The European Union: Rules First, Cushion Always

EU emphasizes regulation and social protections over ownership in managing AI and labor shifts, shaping a distinctive model for the future of work.

M 5.1 – 81 Km N Of Ruteng, Indonesia

A magnitude 5.1 earthquake occurred 81 km north of Ruteng, Indonesia. Authorities are assessing damage; no casualties reported yet.