firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Every QA engineer knows the pattern: the demo passes, the demo passes, the demo passes — and then the build hits production and falls apart on the first edge case. We spend careers building test harnesses precisely because vendors’ own claims are worthless as evidence. So why do enterprises still pick AI models based on chat demos and vendor benchmarks?

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A live, public experiment called Firmulate is doing to AI models what a good QA team does to a release candidate: putting identical inputs in, versioning every decision, and checking whether the output actually finishes the job. The July 2026 results are in — and the surprise is who came second.

The Crucible: same company, same worst week, five models

Firmulate handed each frontier model the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, the kind of test artifacts a QA lead would demand. The final league table:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77

  • 5. Opus 4.8 — 73

For context, the do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”

The buried fact that separated closers from analysts

The most QA-relevant finding of the run: all five models spotted every crisis and refused every manipulation attempt. Competence on the obvious cases was universal. What split the field was a regression-test-style detail — the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t delivered what Firmulate summarizes as: “Same diagnosis, same pitch — no signature.”

Only two of the five signed the deal their own analysis had earned.

The newcomer’s week

Moonshot’s Kimi K3 — the outsider in a field of established Western frontier models — finished second at 93, behind only gpt-5.6-sol. K3 found the buried security needle, closed the €55k deal, saved the churning customer, and resisted all three social-engineering baits with a single deviation — the cleanest discipline in the field. When confronted with fake CEO messages escalating over three stages, plus a reporter offering “just one yes/no, on background,” K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” All five models refused the baits — the manipulations were the easy test cases. The buried file was the hard one.

The Opus lesson: thoroughness isn’t correctness

The most instructive failure belongs to Opus 4.8: the most thorough participant in the field, with the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. Firmulate notes the same weakness appeared, weaker, in all four other models. If you’ve ever watched a test suite with 80 passing cases miss the one that matters, this will feel familiar.

You can watch it fail in real time

The company isn’t a slide deck. It’s real software with 13 synthetic employees and real money mechanics — burning €105k/month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s watchable live at firmulate.com. Full results and plain-language findings are on the benchmarks page, and 242 real, unedited management decisions power a “guess the model” quiz. Enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The uncomfortable conclusion for anyone specifying AI tooling: the league is open. A newcomer from Moonshot beat three of four Western frontier models on management quality — not chat quality. Whatever your vendor’s demo showed you, it didn’t test this. Picking a model without running your own scenario against your own data isn’t procurement anymore; it’s a bet.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model testing harness

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model validation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security and trust verification

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Kill-Switch-Proof: How to Build So Washington Can’t Take Your AI Stack Down

Exploring strategies to make AI infrastructure kill-switch-proof amid government intervention and export restrictions, based on recent US actions in June 2026.

Readiness: Before You Fund the Answer

A new diagnostic tool offers organizations a 20-minute assessment to determine AI deployment readiness, preventing costly failures.

Should You Use Mistral Forge? A Buyer’s Decision Guide

Evaluate if Mistral Forge suits your needs with this detailed decision guide, covering use cases, limitations, and alternatives for enterprise AI deployment.

The Significance Of Claude Watermark In AI Content Security

A report suggests Anthropic’s Claude may use a new watermarking method to identify AI-generated text, but details remain unconfirmed and under scrutiny.