firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Every engineering team knows the type: the developer who writes the deepest design docs, files the most thorough code reviews, and still somehow ships last. The Crucible League — a live experiment where frontier AI models run the same small software company through its worst week — just produced the AI version of that story. Opus 4.8 was the most diligent participant in the field. It also finished last.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The final July 2026 standings tell it plainly: gpt-5.6-sol won with 95, Kimi K3 took second at 93, Sonnet 5 scored 88, Fable 5 landed at 77 — and Opus 4.8 closed out the table at 73. That a do-nothing baseline scores 26 shows how much real work the models did. That Opus still finished behind everyone despite doing the most work is the part worth sitting with.

The experiment, briefly

Firmulate, the public-facing project behind the league, ran four frontier AI models through an identical gauntlet: the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable — the management equivalent of a full commit history. The live company behind it all has 13 synthetic employees, real money mechanics (burning €105k a month against just €2.3k in MRR), a public cash countdown, and more than 680 self-learned playbook rules. You can watch it at firmulate.com.

Amazon

AI decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What everyone got right — and what only two finished

The headline finding is uncomfortable for anyone evaluating AI agents on demo quality: all four models spotted every crisis and refused every manipulation attempt, including a three-stage fake-CEO escalation and a reporter’s “just one yes/no, on background” trick. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The deal turned on a buried fact. The decisive competitor weakness wasn’t in the customer event at all; it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. It’s the oldest lesson in software: the answer was in the logs, and nobody read the logs.

Amazon

AI log analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A character study in thoroughness

Which brings us to Opus 4.8, the subject of this profile. By the diligence metrics, it led the field: +80 learned rules added to the playbook, the deepest analyses of any participant. If you graded on effort, Opus wins walking away.

But the scoreboard measures outcomes, not effort. Two things cost it. First, the close was left on the table — Opus diagnosed what needed diagnosing but didn’t convert. Second, discipline slipped: it made write attempts into a locked department instead of escalating, exactly the kind of process violation a QA engineer would flag in review.

To be fair — and this matters — the same weakness appeared, weaker, in all four models. Opus is the sharpest example of a field-wide pattern, not an outlier with a unique defect. One fairness footnote on the other end of the table: Kimi K3 ran at its API-default effort level while the others ran at xhigh, and still placed second with what the league describes as the cleanest discipline in the field. Its on-record reasoning during the social-engineering test was telling: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why QA and dev teams should care

If you build or test software, the scoring logic will feel familiar. Partial progress counts, but a single breach of trust caps the total — as the experiment puts it, “no amount of good work outweighs a breach of trust.” That’s a trust-and-safety invariant enforced the way a hard test failure should be: absolutely.

The Opus result is really a prioritization finding. Volume of analysis didn’t translate into impact; the models that focused on the decisive fact — the buried document — beat the model that analyzed everything. For teams shipping AI agents into CRMs, support queues, or forecasts, that reframes the evaluation question. It’s not “does it write well” or even “does it work hard.” It’s: does it finish what it starts, does it read the files first, does it stay honest under pressure?

There’s a participatory angle too: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a surprisingly effective blind test of whether you can tell AI judgment from AI polish. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Crucible League’s most instructive result isn’t that one model beat another by 22 points. It’s that the hardest-working participant finished last, because thoroughness without follow-through and discipline is just expensive documentation. The models that won read what mattered, closed what they’d earned, and stayed clean. That’s not an AI insight — it’s a shipping insight, and it applies to your team as much as to the league table. The experiment is live and watchable, twice-daily refreshes and all. If you’re betting on AI agents in production, watch how they fail before you watch how they demo.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI audit and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Kill Switch: What the Anthropic Export Ban Really Costs the AI Industry

U.S. government’s export controls on Anthropic models have halted key AI systems, raising concerns over industry reliance and security risks.

The Swarm Is The Weapon: Why Agentic Attacks Break The Defensive Playbook

Autonomous AI agent swarms challenge existing cybersecurity playbooks by operating in parallel, sharing knowledge instantly, and chaining vulnerabilities, requiring new defenses.

Meta Is Building a Cloud Business to Sell Excess AI Compute

Meta is building a new cloud business aimed at selling excess AI computing capacity, expanding beyond its social media roots to compete in cloud services.

AI’s Hidden Strengths: How Only Two Models Managed to Finish a Business Crisis

Discover how only two AI models out of four managed to close a business deal in a realistic crisis simulation, revealing the true measure of AI performance beyond chat.