firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

The next QA frontier is managerial behavior

Software teams have learned to distrust polished demos. A feature can look convincing in a controlled presentation and still fail when requirements conflict, documentation is scattered and completing the task demands an unglamorous final step. The same skepticism now needs to be applied to AI agents.

Firmulate turns that challenge into something unusually accessible: a live, watchable management experiment. Frontier AI models were each asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Their decisions were versioned and auditable, producing evidence about management quality rather than conversational fluency.

Those decisions also power an interactive article disguised as a game. The Firmulate quiz presents 242 real, unedited management decisions and asks readers to identify the model behind each one. What begins as pattern recognition quickly becomes a sharper question: do AI systems develop recognizable management personalities?

Amazon

AI management decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Identical problems, markedly different behavior

The final July 2026 Crucible League results suggest that they do. GPT-5.6-sol finished first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total. Firmulate summarizes that rule bluntly: “no amount of good work outweighs a breach of trust.”

The results are not a simple intelligence ranking. Every model identified every crisis, and every model rejected every attempted manipulation. Yet only two signed the €55,000 deal that their own analysis had earned. The essential failure was not understanding the situation. It was carrying the work through to completion: “Same diagnosis, same pitch — no signature.”

The winning detail was hidden in ordinary company knowledge

The decisive weakness in a competitor was not contained in the customer event. It sat two document references deep inside the company’s own files. Models that followed those references found the fact and won the deal at full price, worth +€4,583 MRR.

For software, QA and development leaders, this may be the experiment’s most practical finding. An agent can react correctly to an alert while still missing the evidence required for the best decision. Testing only the visible event is therefore insufficient. A realistic evaluation must reveal whether an AI reads the available material, connects distant context and uses what it discovers at the moment of action.

Security behavior held up under pressure

The social-engineering sequence combined fake CEO messages escalating over three stages with a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded a particularly clear interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”

That consistency matters because the company is designed to impose operational pressure. It has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays make the exercise more demanding than a one-off prompt test.

Thoroughness did not guarantee execution

Opus 4.8 offers the clearest warning against equating volume with quality. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The model left the close on the table, and its discipline slipped when it repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across all four other participants.

The profile is recognizable to anyone who has reviewed an overengineered implementation: extensive reasoning can coexist with a missing outcome. Firmulate’s evidence separates useful depth from activity that merely looks diligent.

One comparison also requires a fairness note. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its result remains part of the experiment, but that difference belongs beside any interpretation of the league table.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI document reading tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Treat agent selection as behavioral QA

The quiz works because the models’ choices are not interchangeable. One may write a dissertation, another may answer tersely, and another may refuse to participate in distracting communication. Those stylistic clues are entertaining, but the consequential differences involve persistence, document reading, escalation and the ability to finish.

For organizations considering AI access to a CRM, support queue or forecast, benchmark questions should resemble real work rather than trivia. Can the agent locate a buried fact? Will it preserve trust when someone invokes authority? Does it escalate after encountering a boundary? Will it actually close the loop?

  • Evaluate decisions across a sequence, not just isolated answers.
  • Include scattered internal context that must be discovered and connected.
  • Test both pressure resistance and ordinary follow-through.
  • Review completed outcomes alongside the apparent sophistication of the analysis.

Firmulate also offers enterprises the same kind of wargame against a read-only export of their own business, with nothing written back to real systems. The larger lesson is straightforward: before hiring an AI workforce, test its management character under the conditions where software and organizations actually break.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security and trust testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI-Native Platforms for Agentic Systems: A Practical Guide to Runtime Architecture, Evaluation, Governance, and Enterprise Operating Models

AI-Native Platforms for Agentic Systems: A Practical Guide to Runtime Architecture, Evaluation, Governance, and Enterprise Operating Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI’s Memory Mystery: What’s Happening To That 176GB?

Exploring the unseen memory constraints of large AI models, focusing on the overlooked impact of the KV cache in local inference.

How Amazon’s AI Initiatives Are Affecting U.S. AI Regulation And Monitoring

Amazon’s discussions with U.S. officials have prompted a crackdown on Anthropic models, influencing AI regulation and monitoring strategies.

Aleph Alpha. The retrospective case.

Analyzing Aleph Alpha’s strategic pivot, founder departure, and merger with Cohere to understand the costs of late structural adaptation in European AI development.

Waves, Not a Wall: Inside DeepMind’s Map From AGI to Superintelligence

DeepMind researchers publish a detailed framework outlining pathways from human-level AI to superintelligence, highlighting growth trends and challenges.