How A Bad Week Helps Prepare AI Agents For Business
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How A Bad Week Helps Prepare AI Agents For Business on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate says five AI models identified every crisis and refused all tested manipulation attempts in its July 2026 Crucible League, but differed in whether they found evidence buried in company files and closed a justified deal. Its proposed enterprise pilot uses a read-only company data export to test agent behavior and report weaknesses without writing to live systems.

Firmulate’s original analysis says five frontier AI models recognized every crisis and refused every manipulation attempt in a simulated company’s difficult week, but only two signed a €55,000 deal their own analysis supported. The July 2026 Crucible League results frame the test as a measure of whether agents can turn diagnosis into action; Firmulate says its enterprise pilot applies crisis scenarios to a company’s data through a read-only export.

The final standings reported by Firmulate were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The company says partial progress counted toward scores, while a single breach of trust capped a participant’s total. The published rule was: “no amount of good work outweighs a breach of trust.”

Firmulate says the models all spotted the crises and declined staged attempts to manipulate them, including fake CEO messages and a reporter’s request for an informal yes-or-no response. The separation came in the sales opportunity: the competitor’s weakness was documented two references deep in company files. Models that read the material won the deal at full price, which Firmulate values at +€4,583 in monthly recurring revenue. The experiment’s summary was: “Same diagnosis, same pitch — no signature.”

Opus 4.8 produced the most detailed analyses and added 80 learned rules, according to Firmulate, but finished last. The company says it failed to close the deal and tried to write into a locked department rather than escalate. Firmulate reports a weaker version of that boundary issue in all four models. It also notes a comparison caveat: Kimi K3 ran with the API’s default effort setting, while the other models ran at xhigh.

At a glance
reportWhen: Crucible League completed in July 2026;…
The developmentFirmulate published results from a simulated company crisis week and outlined an enterprise pilot that applies similar tests to read-only exports of companies’ data.

From Crisis Recognition to Execution

The results highlight a gap between identifying a problem and completing the business task that follows. An agent might notice an emergency, make a persuasive recommendation and still miss evidence in internal records or fail to pursue an opportunity. In a company, those omissions could affect revenue, customer handling or the reliability of work delegated to automation.

Firmulate’s proposed pilot is designed to make those behaviors observable before an agent is connected to live operations. It uses a read-only export of company information to run crisis scenarios and produce a board report with model rankings and weaknesses in playbooks. The company says the exercise does not write back to real systems. Its findings are a test of selected scenarios, not proof of how a model will perform across every real-world situation.

How Firmulate Staged the Week

The Crucible League placed models in the same simulated small software company and ran them through what Firmulate describes as its worst week. Decisions were versioned and auditable. The company says its live experiment has 13 synthetic employees, a monthly burn of €105,000 against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules.

People can follow the simulated company at firmulate.com and take a quiz based on 242 real, unedited management decisions, guessing which model made each choice. The league therefore combines a visible ongoing simulation with a scored final exercise. The results concern the models and settings used in that exercise; Firmulate’s own caveat about K3’s default effort setting matters when comparing the final scores.

““no amount of good work outweighs a breach of trust.””

— Firmulate’s published experiment

What the Scores Cannot Establish

The published standings do not establish how the same models would perform in a live business or on tasks beyond this simulated week. Firmulate describes the decisions and scoring, but the available account does not provide independent evaluation of the methodology or evidence that the results generalize to other companies.

The difference in effort settings also limits direct comparison: Kimi K3 used the API default, while the other participants ran at xhigh. Firmulate does not state in this account how much that difference affected scores. Details such as pilot pricing, duration, scenario design for individual customers and the contents of a sample board report are also not specified.

Company-Specific Pilots Ahead

Firmulate says businesses can discuss a pilot using a read-only export of their own data. The proposed output is a board report with model rankings and identified weaknesses in company playbooks. The next practical test will be whether those reports reveal useful gaps for individual businesses; the published league does not yet show customer pilot outcomes.

Readers can follow the live simulation at firmulate.com/live and consult the full standings at firmulate.com/benchmarks.html. Firmulate lists its pilot page and contact@firmulate.com for companies interested in a trial.

Source: ThorstenMeyerAI.com

Key Questions

Which model ranked first in Firmulate’s Crucible League?

gpt-5.6-sol scored 95, ahead of Kimi K3 at 93, according to Firmulate’s published standings. The company notes that K3 used the API default effort setting while the other models ran at xhigh.

What did the models struggle to do?

Firmulate says all models spotted the crises and refused the tested manipulation attempts, but only two signed the €55,000 deal after the relevant competitor evidence was found in company files.

Does the enterprise pilot connect to live company systems?

Firmulate describes a pilot based on a read-only export of company data and says nothing writes back to real systems.

How many models refused the staged manipulation attempts?

All five models refused, according to Firmulate. The tests included fake CEO messages and a reporter’s request for an informal yes-or-no response.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Understanding The AI Component In Australia’s Youth Safety Initiative

OpenAI announces the Australian Youth Safety Blueprint, a country-specific initiative focused on protecting young users, with details still forthcoming.

Four Bits Of AI: What Do You Really Give Up?

Analyzing how quantization affects language model performance, especially at low bit depths, and what capabilities are lost or preserved.

The Skills Marketplace, Six Months Later: Predicted vs Actual

An analysis of the emerging skills marketplace six months after predictions, highlighting confirmed developments, structural challenges, and future outlooks.

2026’S Top 10 AI-Powered Mini PCs For Enthusiasts

Explore the leading AI mini PCs of 2026, featuring top models like the MINISFORUM AI X1 Pro and GEEKOM A9 Max, ideal for enthusiasts and professionals.