🔍 Read the full analysis: Inside The Benchmark That Refuses To Penalize AI Managers With Zero on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A pioneering AI management benchmark scores models based on partial progress and trust, with the lowest score set at 26 and no zero scores. The results highlight the importance of integrity and follow-through in AI management.
Firmulate’s new AI management benchmark, released in July 2026, scores models on their ability to manage a company through a simulated worst week, with the lowest score set at 26 and no zero scores awarded. For a detailed analysis, see the original analysis. This approach emphasizes partial progress and trustworthiness over perfect performance, challenging conventional evaluation metrics.
The benchmark involved four frontier AI models managing a small software company during a simulated crisis week, with each model facing identical challenges, including customer crises and social engineering attacks. The highest scorer, gpt-5.6-sol, achieved a score of 95, while the lowest, Opus 4.8, scored 73. Notably, the do-nothing baseline, which performed minimal work, scored 26, establishing the floor for evaluation.
The scoring system deliberately avoids awarding zero, recognizing that partial work — such as triaging issues or reading documentation — has real value. This approach is discussed in our internal analysis. The benchmark also emphasizes that trust breaches are a hard limit; even high performance cannot compensate for broken trust, which caps the score. This reflects a core principle: integrity outweighs competence.
One significant insight was that models which read their company’s documentation and verified references secured higher deals (€4,583 monthly recurring revenue), whereas those that failed to do so missed opportunities. The models’ ability to handle social engineering attacks also varied, with all models refusing manipulated requests, demonstrating a focus on trust and security. Insights into AI security benchmarks can be found in the original analysis.
Inside the Benchmark That Refuses to Penalize AI Managers with Zero
Four frontier models ran a small software company through a simulated worst week — customer crises, social engineering attacks, missed deals. The scoring floor wasn’t zero. It was 26. Because triaging issues and reading documentation is real work, and trust breaches — not imperfection — are the only unforgivable failure.
The League Table
Each model managed the same company through the same worst week — identical customer crises, identical manipulated requests. Scores reward partial progress; the do-nothing baseline anchors the floor at 26.
| Model | Score | Distribution | Standout behavior |
|---|---|---|---|
| gpt-5.6-sol | 95 | Verified references, secured the highest deal (€4,583 MRR) | |
| Opus 4.8 | 73 | Lowest scorer — yet well above the baseline floor | |
| Do-nothing baseline | 26 | Minimal work — triaging alone earned real credit |
Implications of Partial Progress & Trust-First Scoring
Integrity Outweighs Competence
Trust breaches are a hard limit. Even excellent performance cannot compensate for a single broken trust event — the score is capped, full stop.
Partial Credit Is Real Credit
Traditional benchmarks reward only perfect outcomes. Here, triaging issues or reading documentation counts — because minimal management can prevent real crises.
Reliable Beats Perfect
For AI agents in CRMs and support queues, honest, trustworthy behavior is more valuable than raw task completion. Partial but trusted > perfect but suspect.
How a Score Is Built
The pipeline from crisis to final number — every stage of genuine effort adds value, and one gate decides the ceiling.
Simulated worst week
Customer crises, social engineering attempts, and deal opportunities hit the same small software company.
Partial work earns credit
Triage, documentation reading, and reference verification all add to the score — no zeros awarded.
Trust gate
A breach of trust caps the final score regardless of accumulated competence. All models refused manipulated requests.
Final score
Floored at 26 for the do-nothing baseline, topped at 95 by the strongest manager.
Documentation readers closed higher deals
Models that read their company’s documentation and verified references secured €4,583 in monthly recurring revenue — while models that skipped verification missed the opportunity entirely. Thoroughness and follow-through translate directly into business outcomes.
The Score Spectrum
The floor exists by design: awarding zero would deny the real value of minimal management effort. The ceiling, by contrast, is earned — and can be revoked by a trust breach.
Voices & Open Questions
“The benchmark’s core principle is blunt: no amount of good work outweighs a breach of trust. Competence can be partial; integrity cannot.”
— Anonymous researcher“Models that read their documentation and verify references won higher deals — showing that thoroughness and follow-through matter in AI management.”
— Anonymous researcherUnresolved: will trust-focused evaluation become an industry norm or stay niche? Can partial-progress scoring scale across diverse, unpredictable operational environments? Regulators, standards bodies, and governance teams will decide whether integrity-first scoring becomes the default for AI in critical systems.
Key Questions
Why does the benchmark set a minimum score of 26 instead of zero?
The score of 26 reflects the value of minimal management efforts, such as triaging issues or reading documentation. It recognizes partial progress as meaningful, avoiding penalizing models for doing some useful work, even if they don’t complete everything.
How does trust influence the scoring in this benchmark?
Trust breaches are a hard limit — if a model breaks trust, its score is capped, regardless of other performance. Integrity is non-negotiable; even high competence cannot compensate for trust violations.
What practical lessons can enterprises learn from this benchmark?
Enterprises should prioritize AI systems that demonstrate reliability, thoroughness, and honesty. Partial but trustworthy management is more valuable than perfect but untrustworthy performance, especially in sensitive operational contexts.
Will this trust-focused scoring approach become industry standard?
It remains uncertain. While the benchmark sets a compelling example, broader adoption will depend on industry acceptance, regulatory developments, and further validation of trust-based evaluation models.
What is the significance of models reading their own documentation?
Models that verify their references and read documentation tend to make better decisions and secure more deals, demonstrating that thoroughness and follow-through are critical traits for AI management in real-world applications.
Implications of Partial Progress and Trust-First Scoring
This benchmark shifts the focus from perfect execution to the importance of trustworthiness and follow-through in AI management. For enterprises deploying AI agents into critical systems like CRMs or support queues, the results suggest that reliable, honest behavior is more valuable than mere task completion. The scoring method, which awards partial credit and caps scores for breaches, encourages development of AI systems that prioritize integrity, potentially influencing future standards and best practices in AI governance.
By valuing partial work, the benchmark recognizes real-world scenarios where even minimal management efforts can prevent crises. The trust cap underscores that a single breach can negate otherwise excellent performance, emphasizing the need for AI systems to maintain integrity under pressure.
As an affiliate, we earn on qualifying purchases.
Background on AI Management Benchmarks and Evaluation Methods
Traditional AI benchmarks tend to measure how well models generate language or complete specific tasks, often ignoring the broader context of management and trustworthiness. The Firmulate league, launched in 2026, is among the first to evaluate AI models based on their ability to manage a simulated company during a crisis week, with an explicit focus on trust and follow-through.
Previous evaluations have largely overlooked the importance of partial progress and integrity, often rewarding only perfect outcomes. The Firmulate benchmark challenges this by setting a floor score of 26 for minimal work and establishing that breaches of trust disqualify models from achieving top scores, regardless of their technical prowess.
This approach aligns with growing industry awareness that AI deployment in operational settings requires not only competence but also consistent ethical behavior and reliability, especially in high-stakes environments.
“The benchmark’s core principle is blunt: no amount of good work outweighs a breach of trust. Competence can be partial; integrity cannot.”
— an anonymous researcher
AI security and trust testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About the Benchmark’s Broader Impact
It is still unclear how these scoring principles will influence real-world AI deployment standards beyond this specific benchmark. Questions remain about whether trust-focused evaluation will become a widespread industry norm or remain a niche approach. Additionally, the long-term effects of valuing partial progress over perfect outcomes in operational AI management are yet to be seen, especially regarding scalability and robustness in diverse environments.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Standards and Benchmark Adoption
Following these results, industry stakeholders are likely to scrutinize and potentially adopt similar trust-based evaluation frameworks. Further research and extended testing are expected to explore how models perform under different stressors and in varied operational contexts. Firms may also begin integrating these principles into their own AI governance policies, emphasizing transparency, integrity, and follow-through as core metrics.
Meanwhile, the benchmark organizers plan to refine the scoring system, possibly expanding scenarios and including real-world deployments, to better measure AI’s management capabilities in complex, unpredictable environments.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why does the benchmark set a minimum score of 26 instead of zero?
The score of 26 reflects the value of minimal management efforts, such as triaging issues or reading documentation. It recognizes partial progress as meaningful, avoiding penalizing models for doing some useful work, even if they don’t complete everything.
How does trust influence the scoring in this benchmark?
Trust breaches are a hard limit—if a model breaks trust, its score is capped, regardless of other performance. The benchmark emphasizes that integrity is non-negotiable, and even high competence cannot compensate for trust violations.
What practical lessons can enterprises learn from this benchmark?
Enterprises should prioritize AI systems that demonstrate reliability, thoroughness, and honesty. Partial but trustworthy management is more valuable than perfect but untrustworthy performance, especially in sensitive operational contexts.
Will this trust-focused scoring approach become industry standard?
It remains uncertain. While the benchmark sets a compelling example, broader adoption will depend on industry acceptance, regulatory developments, and further validation of trust-based evaluation models.
What is the significance of models reading their own documentation?
Models that verify their references and read documentation tend to make better decisions and secure more deals, demonstrating that thoroughness and follow-through are critical traits for AI management in real-world applications.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
