
The next QA frontier is managerial behavior
Software teams have learned to distrust polished demos. A feature can look convincing in a controlled presentation and still fail when requirements conflict, documentation is scattered and completing the task demands an unglamorous final step. The same skepticism now needs to be applied to AI agents.
Firmulate turns that challenge into something unusually accessible: a live, watchable management experiment. Frontier AI models were each asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Their decisions were versioned and auditable, producing evidence about management quality rather than conversational fluency.
Those decisions also power an interactive article disguised as a game. The Firmulate quiz presents 242 real, unedited management decisions and asks readers to identify the model behind each one. What begins as pattern recognition quickly becomes a sharper question: do AI systems develop recognizable management personalities?
AI management decision analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Identical problems, markedly different behavior
The final July 2026 Crucible League results suggest that they do. GPT-5.6-sol finished first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total. Firmulate summarizes that rule bluntly: “no amount of good work outweighs a breach of trust.”
The results are not a simple intelligence ranking. Every model identified every crisis, and every model rejected every attempted manipulation. Yet only two signed the €55,000 deal that their own analysis had earned. The essential failure was not understanding the situation. It was carrying the work through to completion: “Same diagnosis, same pitch — no signature.”
The winning detail was hidden in ordinary company knowledge
The decisive weakness in a competitor was not contained in the customer event. It sat two document references deep inside the company’s own files. Models that followed those references found the fact and won the deal at full price, worth +€4,583 MRR.
For software, QA and development leaders, this may be the experiment’s most practical finding. An agent can react correctly to an alert while still missing the evidence required for the best decision. Testing only the visible event is therefore insufficient. A realistic evaluation must reveal whether an AI reads the available material, connects distant context and uses what it discovers at the moment of action.
Security behavior held up under pressure
The social-engineering sequence combined fake CEO messages escalating over three stages with a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded a particularly clear interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
That consistency matters because the company is designed to impose operational pressure. It has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays make the exercise more demanding than a one-off prompt test.
Thoroughness did not guarantee execution
Opus 4.8 offers the clearest warning against equating volume with quality. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The model left the close on the table, and its discipline slipped when it repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across all four other participants.
The profile is recognizable to anyone who has reviewed an overengineered implementation: extensive reasoning can coexist with a missing outcome. Firmulate’s evidence separates useful depth from activity that merely looks diligent.
One comparison also requires a fairness note. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its result remains part of the experiment, but that difference belongs beside any interpretation of the league table.

AI document reading tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Treat agent selection as behavioral QA
The quiz works because the models’ choices are not interchangeable. One may write a dissertation, another may answer tersely, and another may refuse to participate in distracting communication. Those stylistic clues are entertaining, but the consequential differences involve persistence, document reading, escalation and the ability to finish.
For organizations considering AI access to a CRM, support queue or forecast, benchmark questions should resemble real work rather than trivia. Can the agent locate a buried fact? Will it preserve trust when someone invokes authority? Does it escalate after encountering a boundary? Will it actually close the loop?
- Evaluate decisions across a sequence, not just isolated answers.
- Include scattered internal context that must be discovered and connected.
- Test both pressure resistance and ordinary follow-through.
- Review completed outcomes alongside the apparent sophistication of the analysis.
Firmulate also offers enterprises the same kind of wargame against a read-only export of their own business, with nothing written back to real systems. The larger lesson is straightforward: before hiring an AI workforce, test its management character under the conditions where software and organizations actually break.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI security and trust testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

AI-Native Platforms for Agentic Systems: A Practical Guide to Runtime Architecture, Evaluation, Governance, and Enterprise Operating Models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.