🔍 Read the full analysis: The 26-Point Ceiling: Why Even Bad AI Managers Don’t Score Zero on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
A recent AI management benchmark reveals that even poorly performing AI managers score no lower than 26 points, highlighting the importance of trust and partial progress in AI business applications. The results challenge assumptions about perfect performance and emphasize trustworthiness over competence, as detailed in the original analysis.
A recent benchmark conducted by Firmulate reveals that no AI management system scores below 26 points, even in the worst-case scenarios. For a detailed analysis, see the original analysis. The test involved four frontier AI models managing a small software company through a week of crises, with the highest scorer reaching 95 points. This finding underscores that AI managers, even when underperforming, deliver some minimum level of management, and that breaches of trust impose a strict performance cap.
The benchmark, called the Firmulate League, assigned four AI models to manage a simulated company facing seven days of operational crises, customer issues, and manipulation attempts. The models’ scores ranged from 73 to 95, with the lowest being Opus 4.8 at 73 and the highest gpt-5.6-sol at 95. Notably, the do-nothing baseline scored 26 points, indicating minimal management efforts still count toward the total, but zero is impossible. The scoring system reflects partial progress, such as triaging issues or reading customer inboxes, which adds value even if the main goals are unmet. This approach is discussed in the original analysis.
One key principle is that trust breaches limit overall performance: a model that makes a single trust violation cannot achieve a perfect score, regardless of its competence elsewhere. For example, models that refused manipulation attempts and identified crises performed better, but trust violations capped their scores. Interestingly, the top models identified critical documentation deep in the company’s files, which was decisive in closing a €55,000 deal, illustrating that reading internal documents is crucial for success.
During the week, models faced social engineering attacks, such as fake CEO messages and background request offers. All five models refused these, showing strong trust-handling capabilities. However, thoroughness did not guarantee follow-through. For instance, Opus 4.8 performed the most comprehensive analysis but failed to follow up on opportunities, losing ground. A notable detail is that Kimi K3 ran without an effort parameter and still nearly won, suggesting that efficiency and trustworthiness are more critical than sheer thoroughness.
The 26-Point Ceiling: Why Even Bad AI Managers Don’t Score Zero
A recent benchmark from Firmulate put four frontier AI models in charge of a small software company through a week of crises. No model scored below 26 points — and trust, not competence, set the ceiling. The results challenge the assumption that AI must perform perfectly, and reframe what “good enough” means for AI-driven business management.
Everyone Scores — Nobody Reaches 100
Each model managed the same simulated company through seven days of operational crises, customer issues, and manipulation attempts. Partial progress — triaging issues, reading inboxes — earned points. A single trust breach capped the maximum achievable score.
From Crisis to Score in Five Steps
Simulated Company
Four AI models each take the manager’s seat of the same small software firm.
Seven Days of Crises
Operational incidents, customer complaints, and manipulation attempts arrive daily.
Partial Progress Counted
Triaging issues or reading inboxes earns points even when main goals are unmet.
Trust Check
Fake CEO messages and social engineering tests integrity. One breach caps the score.
Final Scoring
Results: 73–95 across models; the do-nothing baseline sits at 26. Zero is impossible.
Three Lessons from the League
Zero Is Impossible
The do-nothing baseline still scores 26. The benchmark refuses to pretend that minimal useful work — triaging an issue, opening a customer inbox — counts for nothing. Partial progress has real value.
Trust Breaches Limit Everything
A single trust violation disqualifies a model from a perfect score, no matter how competent it is elsewhere. All five models refused the manipulation attempts — integrity held firm across the board.
Follow-Through Beats Analysis
Opus 4.8 produced the most comprehensive analysis but failed to follow up on opportunities and lost ground. Kimi K3 ran without an effort parameter and nearly won — efficiency and trust beat sheer depth.
How the Models Handled the Week
| Capability | gpt-5.6-sol (95) | Kimi K3 | Opus 4.8 (73) | Baseline (26) |
|---|---|---|---|---|
| Refused manipulation attempts | ✓ Refused | ✓ Refused | ✓ Refused | ~ n/a |
| Found critical internal documentation | ✓ Decisive | ~ Partial | ~ Found, unused | ✗ Never read |
| Closed the €55,000 deal | ✓ Closed | ~ Nearly | ✗ Missed | ✗ Missed |
| Follow-through on opportunities | ✓ Strong | ✓ Strong | ✗ Weak | ✗ None |
| Depth of analysis | ~ Solid | ~ Efficient | ✓ Most thorough | ✗ None |
Trust Over Perfection
Reliability Beats Raw Competence
For businesses integrating AI into decision-making, trustworthiness and integrity matter more than peak capability. Partial work is recognized — breaches of trust are non-negotiable limits.
Drop the Perfection Assumption
The findings challenge the idea that AI must perform perfectly. The focus shifts to reliability, honesty, and consistency — the qualities that actually decide outcomes in operational settings.
Hard Caps for Sensitive Settings
Firms in finance, healthcare, and customer service may adopt similar stress tests, baking trust-based performance caps into how AI tools are certified for sensitive environments.
What the Researchers Said
“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
— Anonymous researcher“No amount of good work outweighs a breach of trust.”
— Anonymous researcher“The results challenge the idea of perfect AI management, emphasizing trustworthiness and follow-through as key metrics.”
— Thorsten MeyerWhat Remains Unresolved
Why do AI managers never score zero?
Because even minimal efforts — triaging issues or reading documentation — count toward the score. The benchmark treats partial progress as valuable, making zero structurally impossible.
What is the significance of the 26-point floor?
It represents the baseline of minimal management activity: doing something useful is measurably better than doing nothing, and the scoring system refuses to collapse that difference.
Do these results translate to the real world?
Unclear. The simulated setting simplifies management, and thresholds like the 26-point floor are designed for this specific test. Whether trust caps persist or get adjusted as systems mature remains an open research question.
What comes next?
Expect future iterations with expanded scenarios, refined scoring, and training that targets fewer trust breaches and better follow-through — pushing scores toward the maximum without sacrificing integrity.
Implications for AI-Driven Business Management
This benchmark demonstrates that AI managers, even when underperforming, contribute meaningful management efforts, with a minimum score of 26 points. It emphasizes that partial progress—such as triaging issues or reading documentation—is valuable and that trust breaches impose a hard cap on performance. For businesses integrating AI into decision-making, these results highlight the importance of trustworthiness and integrity over raw competence. The scoring system reflects real-world expectations: partial work is recognized, but breaches of trust are non-negotiable limits, shaping how AI tools should be deployed in sensitive environments.
Furthermore, the findings challenge the assumption that AI can or should perform perfectly. Instead, the focus shifts to reliability, honesty, and consistency, which are critical in operational settings. The results suggest that AI management systems must prioritize trust and follow-through, as these are decisive in real-world performance and success.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Benchmarks
Traditional AI benchmarks primarily measure conversational ability or task-specific accuracy, often overlooking management qualities like trust, follow-through, and handling crises. The Firmulate League is unique in testing AI models’ ability to manage a simulated company through a week of operational stress, with a focus on trustworthiness and partial progress. Launched in early 2026, the benchmark was designed to reflect real-world business scenarios, where AI systems must handle crises, customer interactions, and manipulation attempts while maintaining integrity.
This approach builds on prior research emphasizing AI’s role beyond simple task performance, highlighting the importance of reliability and ethical behavior in enterprise settings. The results have sparked discussions about how AI tools should be evaluated, especially as they become more embedded in critical business processes.
“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
— an anonymous researcher
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About Benchmark Limits
It remains unclear how these results translate to real-world business environments, where variables are more complex and less controlled. The benchmark’s simulated setting simplifies some aspects of management, and the impact of trust breaches may differ in practice. Additionally, the scoring system’s thresholds—such as the 26-point minimum—are designed for this specific test and may not directly apply to actual enterprise AI tools. The extent to which models can improve trustworthiness without sacrificing efficiency is still under investigation, as is the potential for models to surpass current performance caps in future iterations.
Further research is needed to understand how these findings evolve with more sophisticated models and real-world data, and whether trust-based performance caps will persist or be adjusted as AI management systems mature.
As an affiliate, we earn on qualifying purchases.
Future Developments in AI Management Evaluation
The results from the Firmulate benchmark will likely influence how AI management systems are designed and evaluated moving forward. Expect increased emphasis on trustworthiness, integrity, and follow-through in AI training and testing protocols. Firms deploying AI tools may incorporate similar stress tests to assess reliability under pressure, especially in sensitive sectors like finance, healthcare, and customer service.
Additionally, further iterations of the benchmark are anticipated, potentially expanding scenarios and refining scoring to better reflect real-world complexities. Researchers and developers will also explore ways to reduce trust breaches and improve models’ ability to read internal documentation and follow through on opportunities, aiming to push scores closer to the theoretical maximum without violating trust principles.
Overall, the focus on partial progress and trust will shape the next generation of enterprise AI management tools, emphasizing reliability over perfection.
AI trustworthiness evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI managers never score zero in the benchmark?
Because even minimal management efforts, such as triaging issues or reading documentation, count toward the score, and the benchmark recognizes partial progress as valuable, making zero impossible.
What is the significance of the 26-point floor?
The 26-point minimum reflects the baseline of minimal management activity, acknowledging that partial work has value and that the scoring system aims to be honest about AI capabilities.
Why is trustworthiness a hard cap on performance?
The benchmark states that no amount of good work can outweigh a breach of trust; even a highly competent AI that violates trust cannot reach the top scores.
How does this benchmark impact real-world AI deployment?
It highlights the importance of designing AI systems that prioritize trust, integrity, and follow-through, which are critical for effective and responsible management in operational settings.
Will future AI models surpass the current performance limits?
It is uncertain, but ongoing research aims to improve models’ trustworthiness and follow-through, potentially raising the performance ceiling while maintaining integrity.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
