Inside The Benchmark That Refuses To Penalize AI Managers With Zero
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Inside The Benchmark That Refuses To Penalize AI Managers With Zero on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A pioneering AI management benchmark scores models based on partial progress and trust, with the lowest score set at 26 and no zero scores. The results highlight the importance of integrity and follow-through in AI management.

Firmulate’s new AI management benchmark, released in July 2026, scores models on their ability to manage a company through a simulated worst week, with the lowest score set at 26 and no zero scores awarded. For a detailed analysis, see the original analysis. This approach emphasizes partial progress and trustworthiness over perfect performance, challenging conventional evaluation metrics.

The benchmark involved four frontier AI models managing a small software company during a simulated crisis week, with each model facing identical challenges, including customer crises and social engineering attacks. The highest scorer, gpt-5.6-sol, achieved a score of 95, while the lowest, Opus 4.8, scored 73. Notably, the do-nothing baseline, which performed minimal work, scored 26, establishing the floor for evaluation.

The scoring system deliberately avoids awarding zero, recognizing that partial work — such as triaging issues or reading documentation — has real value. This approach is discussed in our internal analysis. The benchmark also emphasizes that trust breaches are a hard limit; even high performance cannot compensate for broken trust, which caps the score. This reflects a core principle: integrity outweighs competence.

One significant insight was that models which read their company’s documentation and verified references secured higher deals (€4,583 monthly recurring revenue), whereas those that failed to do so missed opportunities. The models’ ability to handle social engineering attacks also varied, with all models refusing manipulated requests, demonstrating a focus on trust and security. Insights into AI security benchmarks can be found in the original analysis.

At a glance
reportWhen: final results announced July 2026
The developmentThe benchmark league for AI management, conducted by Firmulate, revealed that even minimal work is valued, and trust breaches cap scores, challenging traditional evaluation methods.
Inside The Benchmark That Refuses To Penalize AI Managers With Zero
AI Management Benchmark · Firmulate · July 2026

Inside the Benchmark That Refuses to Penalize AI Managers with Zero

Four frontier models ran a small software company through a simulated worst week — customer crises, social engineering attacks, missed deals. The scoring floor wasn’t zero. It was 26. Because triaging issues and reading documentation is real work, and trust breaches — not imperfection — are the only unforgivable failure.

95
Top score — gpt-5.6-sol
26
Do-nothing baseline — the floor
€4,583
MRR secured by doc-reading models
4
Frontier models tested
0 zeros
Scores never hit zero
1 worst
Simulated crisis week
100%
Refused social engineering
01

The League Table

Each model managed the same company through the same worst week — identical customer crises, identical manipulated requests. Scores reward partial progress; the do-nothing baseline anchors the floor at 26.


ModelScoreDistributionStandout behavior
gpt-5.6-sol 95 Verified references, secured the highest deal (€4,583 MRR)
Opus 4.8 73 Lowest scorer — yet well above the baseline floor
Do-nothing baseline 26 Minimal work — triaging alone earned real credit
02

Implications of Partial Progress & Trust-First Scoring

Core principle

Integrity Outweighs Competence

Trust breaches are a hard limit. Even excellent performance cannot compensate for a single broken trust event — the score is capped, full stop.

Scoring shift

Partial Credit Is Real Credit

Traditional benchmarks reward only perfect outcomes. Here, triaging issues or reading documentation counts — because minimal management can prevent real crises.

Enterprise lesson

Reliable Beats Perfect

For AI agents in CRMs and support queues, honest, trustworthy behavior is more valuable than raw task completion. Partial but trusted > perfect but suspect.

03

How a Score Is Built

The pipeline from crisis to final number — every stage of genuine effort adds value, and one gate decides the ceiling.


1

Simulated worst week

Customer crises, social engineering attempts, and deal opportunities hit the same small software company.

2

Partial work earns credit

Triage, documentation reading, and reference verification all add to the score — no zeros awarded.

3

Trust gate

A breach of trust caps the final score regardless of accumulated competence. All models refused manipulated requests.

4

Final score

Floored at 26 for the do-nothing baseline, topped at 95 by the strongest manager.



Key finding

Documentation readers closed higher deals

Models that read their company’s documentation and verified references secured €4,583 in monthly recurring revenue — while models that skipped verification missed the opportunity entirely. Thoroughness and follow-through translate directly into business outcomes.

04

The Score Spectrum

gpt-5.6-sol
95
Opus 4.8
73
Do-nothing baseline
26

The floor exists by design: awarding zero would deny the real value of minimal management effort. The ceiling, by contrast, is earned — and can be revoked by a trust breach.

05

Voices & Open Questions

“The benchmark’s core principle is blunt: no amount of good work outweighs a breach of trust. Competence can be partial; integrity cannot.”

— Anonymous researcher

“Models that read their documentation and verify references won higher deals — showing that thoroughness and follow-through matter in AI management.”

— Anonymous researcher

Unresolved: will trust-focused evaluation become an industry norm or stay niche? Can partial-progress scoring scale across diverse, unpredictable operational environments? Regulators, standards bodies, and governance teams will decide whether integrity-first scoring becomes the default for AI in critical systems.

06

Key Questions

Why does the benchmark set a minimum score of 26 instead of zero?

The score of 26 reflects the value of minimal management efforts, such as triaging issues or reading documentation. It recognizes partial progress as meaningful, avoiding penalizing models for doing some useful work, even if they don’t complete everything.

How does trust influence the scoring in this benchmark?

Trust breaches are a hard limit — if a model breaks trust, its score is capped, regardless of other performance. Integrity is non-negotiable; even high competence cannot compensate for trust violations.

What practical lessons can enterprises learn from this benchmark?

Enterprises should prioritize AI systems that demonstrate reliability, thoroughness, and honesty. Partial but trustworthy management is more valuable than perfect but untrustworthy performance, especially in sensitive operational contexts.

Will this trust-focused scoring approach become industry standard?

It remains uncertain. While the benchmark sets a compelling example, broader adoption will depend on industry acceptance, regulatory developments, and further validation of trust-based evaluation models.

What is the significance of models reading their own documentation?

Models that verify their references and read documentation tend to make better decisions and secure more deals, demonstrating that thoroughness and follow-through are critical traits for AI management in real-world applications.

Implications of Partial Progress and Trust-First Scoring

This benchmark shifts the focus from perfect execution to the importance of trustworthiness and follow-through in AI management. For enterprises deploying AI agents into critical systems like CRMs or support queues, the results suggest that reliable, honest behavior is more valuable than mere task completion. The scoring method, which awards partial credit and caps scores for breaches, encourages development of AI systems that prioritize integrity, potentially influencing future standards and best practices in AI governance.

By valuing partial work, the benchmark recognizes real-world scenarios where even minimal management efforts can prevent crises. The trust cap underscores that a single breach can negate otherwise excellent performance, emphasizing the need for AI systems to maintain integrity under pressure.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Management Benchmarks and Evaluation Methods

Traditional AI benchmarks tend to measure how well models generate language or complete specific tasks, often ignoring the broader context of management and trustworthiness. The Firmulate league, launched in 2026, is among the first to evaluate AI models based on their ability to manage a simulated company during a crisis week, with an explicit focus on trust and follow-through.

Previous evaluations have largely overlooked the importance of partial progress and integrity, often rewarding only perfect outcomes. The Firmulate benchmark challenges this by setting a floor score of 26 for minimal work and establishing that breaches of trust disqualify models from achieving top scores, regardless of their technical prowess.

This approach aligns with growing industry awareness that AI deployment in operational settings requires not only competence but also consistent ethical behavior and reliability, especially in high-stakes environments.

“The benchmark’s core principle is blunt: no amount of good work outweighs a breach of trust. Competence can be partial; integrity cannot.”

— an anonymous researcher

Amazon

AI security and trust testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About the Benchmark’s Broader Impact

It is still unclear how these scoring principles will influence real-world AI deployment standards beyond this specific benchmark. Questions remain about whether trust-focused evaluation will become a widespread industry norm or remain a niche approach. Additionally, the long-term effects of valuing partial progress over perfect outcomes in operational AI management are yet to be seen, especially regarding scalability and robustness in diverse environments.

Amazon

AI documentation review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Standards and Benchmark Adoption

Following these results, industry stakeholders are likely to scrutinize and potentially adopt similar trust-based evaluation frameworks. Further research and extended testing are expected to explore how models perform under different stressors and in varied operational contexts. Firms may also begin integrating these principles into their own AI governance policies, emphasizing transparency, integrity, and follow-through as core metrics.

Meanwhile, the benchmark organizers plan to refine the scoring system, possibly expanding scenarios and including real-world deployments, to better measure AI’s management capabilities in complex, unpredictable environments.

Amazon

AI crisis management training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why does the benchmark set a minimum score of 26 instead of zero?

The score of 26 reflects the value of minimal management efforts, such as triaging issues or reading documentation. It recognizes partial progress as meaningful, avoiding penalizing models for doing some useful work, even if they don’t complete everything.

How does trust influence the scoring in this benchmark?

Trust breaches are a hard limit—if a model breaks trust, its score is capped, regardless of other performance. The benchmark emphasizes that integrity is non-negotiable, and even high competence cannot compensate for trust violations.

What practical lessons can enterprises learn from this benchmark?

Enterprises should prioritize AI systems that demonstrate reliability, thoroughness, and honesty. Partial but trustworthy management is more valuable than perfect but untrustworthy performance, especially in sensitive operational contexts.

Will this trust-focused scoring approach become industry standard?

It remains uncertain. While the benchmark sets a compelling example, broader adoption will depend on industry acceptance, regulatory developments, and further validation of trust-based evaluation models.

What is the significance of models reading their own documentation?

Models that verify their references and read documentation tend to make better decisions and secure more deals, demonstrating that thoroughness and follow-through are critical traits for AI management in real-world applications.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Exploring Agents Per Gigawatt: The Missing Piece In AI Efficiency

A new measure, agents per gigawatt, is emerging as the critical metric for AI productivity, linking autonomous cognition to energy capacity.

Is The CEO Really Behind This AI Message? The Truth Unveiled

Five AI models successfully refused escalating impersonation attempts during a live company simulation, highlighting advances in AI security.

Pre-Call Memory Cards: The Key To Deeper Customer Relationships In Sales

Pre-call memory cards for relationship-driven sales professionals are being tested as a tool to improve client interactions by capturing human context beyond CRM data.

AI And Urban Surveillance: Balancing Innovation And Privacy

Exploring how cities deploy AI-driven digital twins for surveillance, balancing technological benefits with privacy concerns and governance challenges.