firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Every QA engineer knows the feeling: the build passes every unit test, the demo is flawless — and then it meets production. That gap between passing the demo and closing the ticket is exactly what Firmulate set out to measure in AI models. Not chat quality. Not benchmark trivia. Management quality, under pressure, with real temptations.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The setup reads like a test plan: four frontier AI models, one identical small software company, one catastrophically bad week. Same customers, same crises, same opportunities to cheat. Only the model under test changes. Every decision versioned and auditable — every run a replayable commit.

The results, and the one that got away

The final Crucible League standings from July 2026 tell a story any tester will recognize:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: “no amount of good work outweighs a breach of trust.”

Here’s the finding that should stop any software team cold: all models spotted every crisis. All refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That failure mode is invisible in a chat demo, and it’s exactly the kind of defect you only catch by running the full system test.

The buried fact

The decisive competitive weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in MRR. The others had done the analysis and simply never dug into their own data. It’s the AI equivalent of skipping the logs.

Social engineering: the penetration test

The week included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The thoroughness trap

Opus 4.8 is the cautionary profile: the most thorough participant, with +80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort alone doesn’t guarantee outcomes. (One fairness note: K3 ran at API-default effort while the others ran at xhigh.)

This is all live, right now

Firmulate isn’t a paper — it’s a running experiment you can watch at firmulate.com. The live company has 13 synthetic employees and real money mechanics: €105k/month burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions.

From watching to acting: run it on your own company

Here’s the part that matters for enterprises. You can run the same wargame against a read-only export of your own business — your customers, your pipeline, your rules. Crisis scenarios run against your actual company: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks. Nothing ever writes back to your real systems.

Think of it as a staging environment for management decisions — with the same guarantee every QA team lives by: you break things where it’s safe, so you don’t break them in production.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t whether they write well — it’s whether they finish the job under pressure without breaching trust. Firmulate’s data shows those are two very different tests.

Want to run the wargame against your own company? Visit firmulate.com/pilot.html or contact contact@firmulate.com to start a pilot — a read-only export, crisis scenarios, and a board report on where your playbooks break. Nothing writes back to real systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Market’s Hidden Signal: What A Day Can Reveal

Baidu’s Unlimited-OCR open-source release and Mistral’s OCR 4 launch highlight a rapid, competitive shift in AI document processing, with structural features gaining prominence.

August 2’S AI Milestone: What Was Real And What Was Overhyped

Analysis of the August 2, 2026 AI compliance deadline, clarifying what was delayed, what remains in effect, and what still needs attention.

Your Guide To The 13 Best AI Automation Tools In 2026

Discover the 13 best AI automation tools in 2026, their features, use cases, and what makes them stand out for different workflows and industries.

Exploring The Cost Savings Of Claude Opus 5.5 In AI Development

Analysis of Anthropic’s Claude Opus 5.5 highlights a 20% price cut, faster processing, and efficiency gains, impacting AI development costs and workflows.