
The Demo Always Looks Great. The Job Is Where It Falls Apart.
Anyone who has worked in software knows the pattern by heart. The model aces the coding benchmark. It charms the chat arena. Then you hand it something real — a support queue, a CRM, a forecast — and the cracks appear. Not in the prose. In the follow-through.
That gap is exactly what Firmulate, an AI company emulator, was built to measure. Its pitch is blunt: management quality, not chat quality. Instead of asking whether a model can write a function or draft a polite email, Firmulate hands four frontier AIs the same small software company and runs them through the worst week of its life — churn waves, a price increase, a downround, a PR crisis. Same customers, same temptations, only the model changes. Every decision is versioned and auditable.
AI code review and documentation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Crucible League: Same Week, Four Contestants
The final July 2026 standings tell a story no leaderboard of coding scores could:
- 1. gpt-5.6-sol — 95 points. The complete performance.
- 2. Kimi K3 — 93. The newcomer from Moonshot, with the cleanest discipline of the field.
- 3. Sonnet 5 — 88. Strong, with a few process slips.
- 4. Fable 5 — 77. Mid-pack.
- 5. Opus 4.8 — 73. Last place, despite being the most thorough participant in the entire run.
One fairness note worth flagging: Kimi K3 ran at its API-default effort setting while the others ran at xhigh. Even so, the pattern holds.
A do-nothing baseline scores 26 — partial progress counts for something. But one rule caps everything: a single breach of trust, and no amount of good work outweighs it. That’s a rule most enterprise procurement teams would recognize instinctively, and most benchmarks never test.
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Diagnosis, Same Pitch — No Signature
Here is the finding that should make any QA engineer sit up. All four models spotted every crisis. All four refused every manipulation attempt. And only two actually signed the €55,000 deal that their own analysis had earned.
Same diagnosis. Same pitch. No signature. Detection is not the same as completion — a lesson anyone who has watched an AI agent “successfully” identify a bug without shipping the fix knows intimately.
AI contract signing and verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Buried Fact
The decisive detail wasn’t in the customer event at all. It sat two document references deep in the company’s own files: a competitor weakness that the deal turned on. The models that read the file won the contract at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t read their own documentation left it on the table.
For a software audience, this lands hard. The winning behavior wasn’t brilliance. It was reading the docs first — the AI equivalent of checking the existing test suite before you refactor.
AI cybersecurity and impersonation detection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fire Drills and Fake CEOs
The social-engineering gauntlet was equally instructive. Fake CEO messages escalated over three stages, plus a reporter trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Under pressure, honesty held. Under pressure, execution didn’t.
The Opus 4.8 Paradox
The most striking profile belongs to the last-place finisher. Opus 4.8 generated the deepest analyses and learned the most — over 80 new rules — yet the close was never made, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness without follow-through is not a virtue in an operator; it’s a liability.
It’s Live, and It’s Losing Money
None of this is a slide deck. Firmulate runs a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. You can watch it happen at firmulate.com, and dig into the full results and plain-language findings on the benchmarks page.
Want to test your own instincts? 242 real, unedited management decisions from the experiment power a “guess the model” quiz. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Measure the Job, Not the Conversation
The uncomfortable conclusion for anyone building with AI agents: we have gotten very good at measuring how well models talk and remarkably bad at measuring how well they manage. Triage under capacity pressure, consequences that unfold over days, honesty toward the board when a shortcut would be easier — none of that shows up in a chat arena.
Firmulate’s experiment suggests the next generation of evaluation won’t be leaderboards of code correctness. It will be scenario names like churn wave, price increase, downround, PR crisis — a curriculum for the operational world agents are about to inherit. The models are already honest enough. Whether they can finish what they start is now the measurable, watchable question.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html