
Every QA engineer knows the feeling: the build passes every unit test, the demo is flawless — and then it meets production. That gap between passing the demo and closing the ticket is exactly what Firmulate set out to measure in AI models. Not chat quality. Not benchmark trivia. Management quality, under pressure, with real temptations.
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The setup reads like a test plan: four frontier AI models, one identical small software company, one catastrophically bad week. Same customers, same crises, same opportunities to cheat. Only the model under test changes. Every decision versioned and auditable — every run a replayable commit.
The results, and the one that got away
The final Crucible League standings from July 2026 tell a story any tester will recognize:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: “no amount of good work outweighs a breach of trust.”
Here’s the finding that should stop any software team cold: all models spotted every crisis. All refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That failure mode is invisible in a chat demo, and it’s exactly the kind of defect you only catch by running the full system test.
The buried fact
The decisive competitive weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in MRR. The others had done the analysis and simply never dug into their own data. It’s the AI equivalent of skipping the logs.
Social engineering: the penetration test
The week included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The thoroughness trap
Opus 4.8 is the cautionary profile: the most thorough participant, with +80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort alone doesn’t guarantee outcomes. (One fairness note: K3 ran at API-default effort while the others ran at xhigh.)
This is all live, right now
Firmulate isn’t a paper — it’s a running experiment you can watch at firmulate.com. The live company has 13 synthetic employees and real money mechanics: €105k/month burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions.
From watching to acting: run it on your own company
Here’s the part that matters for enterprises. You can run the same wargame against a read-only export of your own business — your customers, your pipeline, your rules. Crisis scenarios run against your actual company: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks. Nothing ever writes back to your real systems.
Think of it as a staging environment for management decisions — with the same guarantee every QA team lives by: you break things where it’s safe, so you don’t break them in production.

If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t whether they write well — it’s whether they finish the job under pressure without breaching trust. Firmulate’s data shows those are two very different tests.
Want to run the wargame against your own company? Visit firmulate.com/pilot.html or contact contact@firmulate.com to start a pilot — a read-only export, crisis scenarios, and a board report on where your playbooks break. Nothing writes back to real systems.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
