How The Post-Demo Leaderboard Shapes AI Innovation
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A live experiment by Firmulate tests AI models in managing a small company during its worst week. Results show management skills, trustworthiness, and decision quality are key, not just chat or coding ability. This could redefine how AI effectiveness is measured.

In a groundbreaking live experiment, Firmulate has tested AI models in a simulated management role within a small software company during its most challenging week. The results highlight that management quality and trustworthiness are crucial metrics, surpassing traditional chat or coding benchmarks, and could reshape how AI’s business capabilities are evaluated. This approach aligns with emerging standards discussed in industry analyses.

The experiment, called the Crucible League, involved five leading AI models competing to manage a company facing multiple crises, including customer churn, PR issues, and financial pressures. For more context, see the original analysis. The models were scored based on their ability to diagnose problems, communicate decisions, and maintain trust. GPT-5.6-SOL topped the leaderboard with a score of 95, while Opus 4.8 lagged at 73. Notably, all models successfully identified crises and rejected manipulation attempts, but only two managed to close deals worth €55,000, illustrating a gap between diagnosis and action.

One key finding was that models could sound confident and informed but fail to retrieve critical facts buried deep in company files, leading to missed opportunities. For a deeper dive into this issue, see the original analysis. For example, a model that read the relevant document references would win a €4,583 MRR deal, while others missed the crucial detail, underscoring that effective management depends on accurate information retrieval, not just eloquence.

Additionally, the experiment tested models against social engineering attempts, such as fake CEO messages and impersonation tricks. All five models refused to comply, demonstrating strong adherence to safety protocols. However, even the most thorough model struggled with executing management tasks effectively, revealing that effort and detailed analysis do not necessarily translate into successful outcomes.

At a glance
reportWhen: ongoing, with final results published i…
The developmentFirmulate’s live management experiment ranks AI models based on their ability to handle a simulated company’s crises, revealing new insights into AI’s role in business management.

Implications for AI Evaluation and Business Use

This experiment shifts the focus from traditional AI benchmarks—such as chat quality or coding prowess—to management skills, trustworthiness, and consequence management. It suggests that future AI evaluation should incorporate real-world decision-making, including handling crises, maintaining trust, and completing tasks reliably. For businesses, this means that selecting AI tools will require assessing their ability to read organizational context, prioritize actions, and uphold integrity under pressure, rather than just their conversational or technical capabilities.

The findings emphasize that AI models must be tested in scenarios that mirror operational environments, where trust and accountability are critical. As AI begins to take on more managerial roles, understanding its capacity to manage consequences without compromising organizational integrity becomes essential. This could influence how AI vendors develop and market their products and how companies integrate these tools into their workflows.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Testing and Firmulate’s Approach

Traditional AI benchmarks focus on isolated skills: coding competitions, chat arenas, or language understanding tests. These, however, do not capture how AI performs in managing real-world tasks that involve multiple steps, trust, and consequences. Recognizing this gap, Firmulate launched a live experiment in July 2026, where AI models managed a small company facing crises similar to those in actual business operations.

The company’s setup includes 13 synthetic employees, real financial mechanics, and a set of 680+ self-learned rules that govern management behavior. The experiment simulates scenarios such as customer churn, PR crises, pricing decisions, and internal escalation, providing a comprehensive testbed for evaluating AI in operational roles. This approach aims to measure not just what models say but what they do—how they diagnose, decide, communicate, and follow through.

Prior to this, most benchmarks did not account for the complexities of trust, consequence management, or organizational context, making the Firmulate experiment a significant step toward more realistic AI evaluation.

“The real test of AI in management isn’t just answering well—it’s managing consequences, maintaining trust, and completing tasks reliably under pressure.”

— Thorsten Meyer, founder of Firmulate

Amazon

business AI simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About Long-Term AI Management Capabilities

While the experiment provides valuable insights, it remains unclear how these models will perform in extended, real-world operational settings beyond the controlled simulation. Questions also persist about how models handle evolving crises over longer periods and whether their trustworthiness can be maintained consistently. Additionally, the impact of different organizational contexts and the scalability of such evaluations are still being explored.

Amazon

AI crisis management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Testing and Adoption

Following the initial results, firms and researchers are expected to develop more comprehensive benchmarks that incorporate longer-term management tasks, real-time decision-making, and trust assessments. Companies considering AI for operational roles will likely demand deeper validation, including live pilot programs and scenario testing tailored to their specific environments. The industry may also see increased emphasis on AI models’ ability to read organizational context, escalate appropriately, and uphold integrity over time.

Further research will aim to refine these evaluation methods, potentially leading to standardized management benchmarks that go beyond current chat and code tests, shaping the future of AI adoption in business management.

Amazon

organizational trust AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How does the Firmulate experiment differ from traditional AI benchmarks?

It tests AI models in managing a simulated company during crises, focusing on decision-making, trust, and consequence management rather than just chat quality or coding skills.

What are the key findings from the July 2026 leaderboard?

Models can diagnose crises and refuse manipulation but often fail to execute decisions effectively or retrieve critical facts, affecting real-world outcomes.

Why is trustworthiness important in AI management?

Trustworthiness ensures AI models act reliably, maintain organizational integrity, and avoid compromising decisions, especially under pressure or manipulation attempts.

Will this change how companies evaluate AI tools?

Yes, companies are likely to adopt more operational and consequence-based benchmarks, emphasizing real-world decision-making and trust over traditional chat or coding tests.

What are the limitations of the current experiment?

It is a controlled simulation, so long-term performance, adaptability to evolving crises, and scalability to larger organizations remain untested.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Regulatory Vacuum.

Google discloses a zero-day AI vulnerability on May 11, 2026, exposing a lack of regulatory frameworks for AI-driven threats, raising urgent policy concerns.

Mac vs GPU Tower for Local LLMs: The Heat-and-Noise Tradeoff

A detailed comparison of Mac Studio with Apple Silicon and GPU towers for running local large language models, focusing on heat, noise, and performance tradeoffs.

Sovereignty Is a Pipe, Not a Passport

Exploring how data sovereignty depends on legal jurisdiction over data handlers, not server location, with insights from Mistral’s model deployment.

Anthropic’s Safety Story Has Become a Power Story

Anthropic reports its models are increasingly autonomous, with over 80% of code now generated by AI, raising questions about AI self-improvement and governance.