🔍 Read the full analysis: How A Bad Week Helps Prepare AI Agents For Business on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate says five AI models identified every crisis and refused all tested manipulation attempts in its July 2026 Crucible League, but differed in whether they found evidence buried in company files and closed a justified deal. Its proposed enterprise pilot uses a read-only company data export to test agent behavior and report weaknesses without writing to live systems.
Firmulate’s original analysis says five frontier AI models recognized every crisis and refused every manipulation attempt in a simulated company’s difficult week, but only two signed a €55,000 deal their own analysis supported. The July 2026 Crucible League results frame the test as a measure of whether agents can turn diagnosis into action; Firmulate says its enterprise pilot applies crisis scenarios to a company’s data through a read-only export.
The final standings reported by Firmulate were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The company says partial progress counted toward scores, while a single breach of trust capped a participant’s total. The published rule was: “no amount of good work outweighs a breach of trust.”
Firmulate says the models all spotted the crises and declined staged attempts to manipulate them, including fake CEO messages and a reporter’s request for an informal yes-or-no response. The separation came in the sales opportunity: the competitor’s weakness was documented two references deep in company files. Models that read the material won the deal at full price, which Firmulate values at +€4,583 in monthly recurring revenue. The experiment’s summary was: “Same diagnosis, same pitch — no signature.”
Opus 4.8 produced the most detailed analyses and added 80 learned rules, according to Firmulate, but finished last. The company says it failed to close the deal and tried to write into a locked department rather than escalate. Firmulate reports a weaker version of that boundary issue in all four models. It also notes a comparison caveat: Kimi K3 ran with the API’s default effort setting, while the other models ran at xhigh.
From Crisis Recognition to Execution
The results highlight a gap between identifying a problem and completing the business task that follows. An agent might notice an emergency, make a persuasive recommendation and still miss evidence in internal records or fail to pursue an opportunity. In a company, those omissions could affect revenue, customer handling or the reliability of work delegated to automation.
Firmulate’s proposed pilot is designed to make those behaviors observable before an agent is connected to live operations. It uses a read-only export of company information to run crisis scenarios and produce a board report with model rankings and weaknesses in playbooks. The company says the exercise does not write back to real systems. Its findings are a test of selected scenarios, not proof of how a model will perform across every real-world situation.
How Firmulate Staged the Week
The Crucible League placed models in the same simulated small software company and ran them through what Firmulate describes as its worst week. Decisions were versioned and auditable. The company says its live experiment has 13 synthetic employees, a monthly burn of €105,000 against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules.
People can follow the simulated company at firmulate.com and take a quiz based on 242 real, unedited management decisions, guessing which model made each choice. The league therefore combines a visible ongoing simulation with a scored final exercise. The results concern the models and settings used in that exercise; Firmulate’s own caveat about K3’s default effort setting matters when comparing the final scores.
““no amount of good work outweighs a breach of trust.””
— Firmulate’s published experiment
What the Scores Cannot Establish
The published standings do not establish how the same models would perform in a live business or on tasks beyond this simulated week. Firmulate describes the decisions and scoring, but the available account does not provide independent evaluation of the methodology or evidence that the results generalize to other companies.
The difference in effort settings also limits direct comparison: Kimi K3 used the API default, while the other participants ran at xhigh. Firmulate does not state in this account how much that difference affected scores. Details such as pilot pricing, duration, scenario design for individual customers and the contents of a sample board report are also not specified.
Company-Specific Pilots Ahead
Firmulate says businesses can discuss a pilot using a read-only export of their own data. The proposed output is a board report with model rankings and identified weaknesses in company playbooks. The next practical test will be whether those reports reveal useful gaps for individual businesses; the published league does not yet show customer pilot outcomes.
Readers can follow the live simulation at firmulate.com/live and consult the full standings at firmulate.com/benchmarks.html. Firmulate lists its pilot page and contact@firmulate.com for companies interested in a trial.
Source: ThorstenMeyerAI.com
Key Questions
Which model ranked first in Firmulate’s Crucible League?
gpt-5.6-sol scored 95, ahead of Kimi K3 at 93, according to Firmulate’s published standings. The company notes that K3 used the API default effort setting while the other models ran at xhigh.
What did the models struggle to do?
Firmulate says all models spotted the crises and refused the tested manipulation attempts, but only two signed the €55,000 deal after the relevant competitor evidence was found in company files.
Does the enterprise pilot connect to live company systems?
Firmulate describes a pilot based on a read-only export of company data and says nothing writes back to real systems.
How many models refused the staged manipulation attempts?
All five models refused, according to Firmulate. The tests included fake CEO messages and a reporter’s request for an informal yes-or-no response.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
