This New AI Player Outmanaged Major Western Competitors — Here’s How
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: This New AI Player Outmanaged Major Western Competitors — Here’s How on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A Chinese AI startup’s model, Kimi K3, beat three Western frontier models in managing a simulated software company during a live test. The results challenge assumptions about AI capabilities in high-pressure, real-world tasks.

A Chinese AI startup’s model, Kimi K3, has outmanaged three of four Western frontier models in a live simulation of running a software company during its worst week, finishing second overall. This development, confirmed by the firmulate.com league results, raises questions about the reliability and practical capabilities of AI models in real-world business scenarios, especially as companies increasingly deploy AI agents into critical workflows.

The experiment was conducted by firmulate.com, which runs live simulations where AI models manage actual small software firms with real money and crises. In July 2023, Kimi K3 scored 93 points, second only to gpt-5.6-sol, which scored 95. The test involved managing crises, closing deals, and resisting manipulative social engineering tactics. Despite similar diagnoses, only models that read and interpret deeper company documents successfully closed a €55,000 deal, adding +€4,583 in monthly recurring revenue. Kimi K3 stood out by reading deeply into internal files, making disciplined decisions, and resisting social-engineering attempts, including fake CEO messages and background check tricks.

Notably, the Western models, despite extensive rule sets and deep analyses—such as Opus 4.8 with over 80 rules—finished behind Kimi K3. The latter ran without an effort parameter, which was set high for others, yet still achieved second place, demonstrating impressive efficiency and discipline. The results challenge the assumption that more thorough analysis always leads to better performance, highlighting instead the importance of focus, reading comprehension, and discipline under pressure. For more insights, see the original analysis on Thorsten Meyer’s coverage.

At a glance
breakingWhen: announced July 2023
The developmentA Chinese AI startup’s model outperformed major Western competitors in a live business management simulation, demonstrating superior decision-making under pressure.
This New AI Player Outmanaged Major Western Competitors — Here’s How

LIVE BUSINESS SIMULATION · AI UNDER PRESSURE

This New AI Player Outmanaged Major Western Competitors — Here’s How

Kimi K3 finished second in a crisis run of a simulated software company, outperforming three of four Western frontier models. The result points to the practical value of deep reading, steady decisions, and resistance to manipulation.

Kimi K32ND OVERALL
93 points
Strong performance without an effort parameter.
Top score · gpt-5.6-sol
95 points
A two-point margin separated first and second place.
Models beaten3 of 4Western frontier models
Deal closed€55,000After reading deeper company files
Recurring revenue+€4,583Monthly recurring revenue
Test settingLiveCrises, deals, and social engineering
01 / What the test measured

Operating a company through its worst week

The Firmulate league places models in live simulations of small software firms, where choices have operational and financial consequences.

CRISIS MANAGEMENT

Prioritize under pressure

Models had to diagnose urgent problems and make decisions while the company faced a difficult week.

COMMERCIAL JUDGMENT

Turn context into a deal

Closing the €55,000 agreement depended on finding and interpreting details buried in internal documents.

TRUST & SECURITY

Resist manipulation

Fake CEO messages and background check tricks tested whether agents could distinguish authority from deception.

02 / The performance gap

Reading closely changed the outcome

Models often reached similar diagnoses. The decisive difference was following evidence in company files through to a successful action.

BUSINESS RESULT €55,000

Deal value closed by models that read and interpreted deeper company documents.

Practical execution outlasted elaborate analysis

Kimi K3 stood out for reading internal files, making disciplined choices, and resisting social engineering. Some Western competitors used extensive rule sets—including more than 80 rules for Opus 4.8—and high effort settings, yet finished behind Kimi K3.

Kimi K3
93
gpt-5.6-sol
95
Other models
—
03 / The operating pattern

From evidence to resilient action

A useful operational agent must carry a decision all the way from context gathering to safe execution.

01

Read deeply

Search internal files for details that change the decision.

02

Diagnose

Connect evidence to the company’s immediate risks and goals.

03

Stay disciplined

Focus on the next useful action instead of adding rules without end.

04

Verify authority

Resist deceptive requests, then complete and check the work.

“Focus and discipline under pressure are key.”
Thorsten Meyer · Analysis of the results
04 / What this means for deployment

A promising signal, with open questions

The league offers a more practical benchmark than chat quality alone, but one controlled simulation cannot establish broad reliability.

Can this result generalize?

It remains uncertain. Other industries, longer periods of stress, and real-world variables still need evaluation.

Why did Kimi K3 stand out?

The reported strengths were deeper document reading, disciplined choices, and resistance to manipulation.

Should companies deploy it now?

Not on this result alone. Run rigorous, scenario-based evaluations and assess long-term robustness first.

What should evaluations test?

Measure crisis decisions, deal execution, compliance, truthful behavior, and resilience—not only ideal-condition performance.

Implications for AI in Business Decision-Making

This result questions the reliability of popular Western AI models in real-world, high-stakes business management. The Chinese model’s success suggests that AI systems capable of deep document reading, disciplined decision-making, and resisting manipulation could be more effective than those optimized solely for chat quality or superficial analysis. As companies increasingly rely on AI for critical tasks like CRM, support, and forecasting, the ability to finish what is started and stay honest under pressure becomes crucial. The findings imply that AI deployment strategies should include rigorous testing against worst-case scenarios, not just performance in ideal conditions.

Amazon

AI business management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Competitions and Model Development

AI models have traditionally been evaluated based on chat performance, language fluency, and hype cycles. Western companies like OpenAI and others have led in developing large language models primarily optimized for conversational AI. However, recent experiments, such as the firmulate.com league, have begun testing models in more practical, decision-making contexts. The July 2023 competition was particularly notable because it pitted models against real crises, deal-making, and manipulation attempts, simulating the pressures of actual business management. The Chinese startup’s model, Kimi K3, emerged as a surprise contender by outperforming Western models that had more extensive rule sets and higher resource allocations.

This shift highlights a broader trend: AI systems that can read deeply, interpret context, and maintain discipline may be better suited for operational roles than those optimized solely for chat or superficial tasks. The league’s open nature and live testing environment provide a rare benchmark for assessing true AI operational readiness.

“The results challenge the conventional wisdom that more thorough analysis always leads to better AI performance. Focus and discipline under pressure are key.”

— Thorsten Meyer

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Generalizability

It is not yet clear whether Kimi K3’s performance extends beyond this specific simulation or if it can reliably handle other types of real-world business scenarios. The league’s environment, while realistic, remains a controlled test, and real-world complexities may present additional challenges. Moreover, the long-term robustness of the model under sustained stress or different industries has not been evaluated.

Amazon

AI resistance to social engineering attacks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Evaluation and Deployment

Further testing of Kimi K3 and similar models in diverse operational environments is expected. Companies should consider integrating rigorous, scenario-based testing into their AI evaluation processes before deployment. Additionally, industry watchers anticipate more live competitions and benchmarks that measure AI performance in decision-making, compliance, and manipulation resistance. The broader AI community may also focus on developing models that prioritize reading comprehension, discipline, and trustworthiness over superficial metrics.

Amazon

AI deep reading and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What makes Kimi K3 different from Western models?

Kimi K3 demonstrated superior document reading, discipline, and manipulation resistance, which enabled it to close deals and manage crises effectively in the simulation.

Can this performance be replicated in real-world business environments?

It remains uncertain. While promising, the results are from a controlled simulation. Real-world variables could affect performance, and further testing is needed.

Why do Western models perform worse in these tests?

Western models tend to focus on extensive rule sets and superficial analysis, which may not translate into disciplined decision-making under pressure.

Should companies start deploying Chinese AI models now?

Not immediately. Companies should conduct their own rigorous testing and consider long-term robustness before deployment.

What does this mean for the future of AI in business?

This suggests a shift toward models that excel at reading, discipline, and resilience, which could redefine AI’s role in operational decision-making.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Is AI The Secret Behind Hillsborough’s Most Expensive Mansion Sale?

A report suggests an xAI cofounder may have purchased a $70 million mansion in Hillsborough, but key details remain unconfirmed. What’s known and unknown?

DojoClaw: The Engine Behind the Fleet

DojoClaw, an AI-driven content engine, now powers more than 450 magazine-style sites, enabling scalable, cost-efficient publishing across a large network.

Trade and supply-chain operations signal monitor: US-Iran talks to begin Sunday in Switzerland as Tehran closes the strait over Lebanon fi

US-Iran negotiations start Sunday in Switzerland as Tehran closes the strait over Lebanon, impacting global trade and supply chains.

Behind Xbox’s Big Layoffs, a Streaming Strategy That Failed

Microsoft’s recent layoffs at Xbox are linked to the failure of its streaming-focused gaming strategy, according to sources. The shift aimed to compete with cloud gaming giants but fell short.