🔍 Read the full analysis: Why The Worst AI Managers Still Achieve A 26 In Industry Benchmarks on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
A recent AI management benchmark shows the lowest-performing AI managers still score 26 points, emphasizing that partial work and trust are valued over perfect but untrustworthy performance. The results challenge assumptions about AI competence and integrity.
The July 2026 results of the Firmulate benchmark league show that the lowest-scoring AI managers, representing the worst performance in managing a small software company during a simulated crisis week, still achieved a score of 26 points. This is significant because it demonstrates that even ineffective AI managers receive a baseline score above zero, highlighting that partial progress and basic trustworthiness are recognized and valued in industry assessments. The findings challenge common perceptions about AI management quality and raise questions about what benchmarks truly measure, as detailed in the original analysis.
The benchmark involved four frontier AI models managing the same company through seven days of crises, customer interactions, and trust tests. For more on industry benchmarks, see Evaluating Apple’s SpeechAnalyzer API. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. The do-nothing baseline, intentionally designed to do almost nothing, scored 26, indicating that even minimal engagement is recognized as partial management effort. The scoring system emphasizes that breaches of trust, even if performance is otherwise strong, result in significant penalties, with a strict cap preventing perfect scores. This approach aligns with the insights from the original analysis. The benchmark’s design aims to reflect real-world management challenges, including handling crises, reading documentation, and resisting social engineering attempts.
One key insight is that models which read their own documentation and verify information performed better, especially in closing deals and avoiding manipulation. Conversely, models that failed to consult relevant files or follow through on tasks scored lower, regardless of their superficial thoroughness. The results suggest that trustworthiness and task completion are more critical than raw conversational ability, aligning with real business priorities.
Why The Worst AI Managers Still Achieve A 26 In Industry Benchmarks
The July 2026 Firmulate results show the lowest-performing AI managers still scored 26 of 100 points. Partial work and basic trustworthiness are rewarded over perfect but unreliable performance — challenging common assumptions about what AI competence really means.
Partial Effort Beats Zero — But Trust Caps Perfection
Four frontier AI models managed the same small software company through seven days of simulated crises, customer interactions, and trust tests. Even the baseline agent — designed to do almost nothing — earned 26 points, proving minimal engagement is recognized as partial management effort.
Seven Days, One Company, Constant Pressure
The Firmulate design mirrors real-world management challenges: reading documentation, closing deals, triaging incidents, and resisting social engineering attempts.
Simulated Crisis Week
Each model manages the same company through seven days of escalating operational crises.
Read & Verify
Models that consult their own documentation and verify claims close more deals and resist manipulation.
Trust Tests
Social engineering probes measure integrity; breaches trigger heavy penalties regardless of output.
Scored Outcome
Partial progress earns points; distrust caps results at the top of the scale.
Models that read documentation and verified information outperformed those with superficial thoroughness. Models that skipped relevant files or failed to follow through scored lower — regardless of conversational polish.
Trust and Follow-Through Outweigh Raw Ability
The scoring system penalizes breaches of trust heavily, while crediting basic task completion — a shift from traditional language-focused benchmarks toward operational integrity.
| Behavior | Recognized by Score | Real-World Analogy | Impact |
|---|---|---|---|
| Reading own documentation | ✓ Yes | Manager reviews the handbook before acting | Higher |
| Verifying customer claims | ✓ Yes | Double-checking before committing | Higher |
| Completing tasks partially | ~ Partial | Minimum viable management | Baseline 26 |
| Fluent conversation only | ✗ No | Charisma without execution | Lower |
| Breach of trust | ✗ Penalized | Broken promise to a customer | Capped score |
What This Means for Businesses and Benchmarks
The results challenge the narrative that AI must be flawless to be useful. Deploying AI agents successfully means prioritizing task completion and integrity over conversational skill — and recognizing that current industry benchmarks may overvalue superficial capabilities while underestimating trust and follow-through.
Reliability Is Paramount
When AI manages critical functions like customer support or sales, integrity under pressure matters more than polish.
Measure What Counts
Expect refined tests of trust, follow-through, and documentation reading — plus more diverse scenarios and real-world deployment validation.
Generalization Unproven
How representative simulated crises are of real management across industries — and how partial scores map to deployment success — remains unvalidated.
Five Answers Buyers Should Know
Why do even the worst AI managers score above zero?
The scoring recognizes partial management efforts — reading documentation, triaging crises, maintaining basic operations — even when tasks go unfinished or trust is breached.
What does a score of 26 really represent?
Minimal management activity: doing just enough to avoid being completely ineffective, reflecting real-world minimum expectations.
Why is untrustworthiness so heavily penalized?
In operational management, trust breaches can have severe consequences; even high-performing models are penalized for violations.
Could a model with a perfect score be untrustworthy?
Yes — which is why a perfect score is treated as suspicious. Trust breaches deliberately prevent reaching full marks.
How should companies interpret these results?
Favor models that complete tasks consistently, read relevant documentation, and maintain integrity under pressure over conversational brilliance.
Implications of Partial Management and Trust in AI
The benchmark results underscore that in AI management, trust and task completion are more influential than perfect performance. Even the worst AI managers score above zero because they perform some minimal management tasks, reflecting real-world expectations where partial progress is valuable. This challenges the narrative that AI must be flawless to be useful; instead, consistent, trustworthy actions matter more. For businesses, this means that deploying AI agents requires a focus on their ability to finish tasks and maintain integrity, not just their conversational skills. The strict penalty for breaches of trust highlights that reliability is paramount, especially when AI manages critical functions like customer support or sales.
Moreover, the results suggest that current industry benchmarks may overvalue superficial capabilities while underestimating the importance of foundational trust and follow-through. As AI systems become more embedded in operational workflows, understanding these priorities will be crucial for effective deployment and risk management.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Benchmarks and Testing
The Firmulate benchmark league was launched to evaluate AI managers based on their ability to handle real-world management tasks during a simulated crisis week. Unlike traditional benchmarks that focus solely on language generation or problem-solving, this test measures how well AI models manage a company’s day-to-day operations, including crisis response, documentation reading, and trustworthiness under social engineering attacks. The July 2026 results build on earlier efforts to quantify AI management effectiveness, emphasizing that partial, trustworthy work is more valuable than perfect but untrustworthy performance. The benchmark’s design reflects growing industry concerns about AI’s role in operational decision-making and the importance of integrity alongside competence.
“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
— an anonymous researcher
Unresolved Questions About Benchmark Interpretation
It is not yet clear how representative the benchmark’s simulated crises are of real-world management challenges across different industries. Additionally, the extent to which partial scores reflect practical AI deployment success remains to be validated in real operational environments. The impact of model-specific configurations, such as the absence of effort parameters in some models, on overall performance and trustworthiness also requires further investigation. The long-term implications of these findings for AI regulation and industry standards are still evolving, and ongoing research is needed to confirm whether these results generalize beyond the specific benchmark setup.
Future Directions for AI Management Evaluation
Researchers and industry practitioners will likely refine benchmarks to better capture the importance of trust, follow-through, and documentation reading. Expect further experiments that include more diverse scenarios, broader industry contexts, and real-world deployment tests. Companies considering AI management tools should prioritize models that demonstrate consistent task completion and integrity, even if their conversational skills are average. The ongoing development of standards and regulations may incorporate these insights to ensure AI systems are reliable and trustworthy in critical operational roles. Additionally, more transparency around how scores are derived will help organizations make informed choices about AI adoption.
Key Questions
Why do even the worst AI managers score above zero?
Because the scoring system recognizes partial management efforts, such as reading documentation, triaging crises, and maintaining basic operations, even if the AI fails to complete all tasks or breaches trust at some point.
What does a score of 26 really represent?
The score indicates minimal management activity, akin to doing just enough to avoid being completely ineffective, with the baseline designed to reflect real-world minimum expectations.
Why is trustworthiness so heavily penalized in the benchmark?
The benchmark emphasizes that in operational management, breaches of trust can have severe consequences, so even high-performing models are penalized for any trust violations, aligning with real-world risk considerations.
Could a model with a perfect score be untrustworthy?
According to the benchmark’s design, a perfect score is suspicious because it might indicate unmeasured or untrustworthy behavior; trust breaches prevent reaching full marks.
How should companies interpret these results for their AI investments?
Focus on models that demonstrate consistent task completion, reading relevant documentation, and maintaining integrity under pressure, rather than just conversational skill or superficial performance.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
