🔍 Read the full analysis: This New AI Player Outmanaged Major Western Competitors — Here’s How on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A Chinese AI startup’s model, Kimi K3, beat three Western frontier models in managing a simulated software company during a live test. The results challenge assumptions about AI capabilities in high-pressure, real-world tasks.
A Chinese AI startup’s model, Kimi K3, has outmanaged three of four Western frontier models in a live simulation of running a software company during its worst week, finishing second overall. This development, confirmed by the firmulate.com league results, raises questions about the reliability and practical capabilities of AI models in real-world business scenarios, especially as companies increasingly deploy AI agents into critical workflows.
The experiment was conducted by firmulate.com, which runs live simulations where AI models manage actual small software firms with real money and crises. In July 2023, Kimi K3 scored 93 points, second only to gpt-5.6-sol, which scored 95. The test involved managing crises, closing deals, and resisting manipulative social engineering tactics. Despite similar diagnoses, only models that read and interpret deeper company documents successfully closed a €55,000 deal, adding +€4,583 in monthly recurring revenue. Kimi K3 stood out by reading deeply into internal files, making disciplined decisions, and resisting social-engineering attempts, including fake CEO messages and background check tricks.
Notably, the Western models, despite extensive rule sets and deep analyses—such as Opus 4.8 with over 80 rules—finished behind Kimi K3. The latter ran without an effort parameter, which was set high for others, yet still achieved second place, demonstrating impressive efficiency and discipline. The results challenge the assumption that more thorough analysis always leads to better performance, highlighting instead the importance of focus, reading comprehension, and discipline under pressure. For more insights, see the original analysis on Thorsten Meyer’s coverage.
LIVE BUSINESS SIMULATION · AI UNDER PRESSURE
This New AI Player Outmanaged Major Western Competitors — Here’s How
Kimi K3 finished second in a crisis run of a simulated software company, outperforming three of four Western frontier models. The result points to the practical value of deep reading, steady decisions, and resistance to manipulation.
Operating a company through its worst week
The Firmulate league places models in live simulations of small software firms, where choices have operational and financial consequences.
Prioritize under pressure
Models had to diagnose urgent problems and make decisions while the company faced a difficult week.
Turn context into a deal
Closing the €55,000 agreement depended on finding and interpreting details buried in internal documents.
Resist manipulation
Fake CEO messages and background check tricks tested whether agents could distinguish authority from deception.
Reading closely changed the outcome
Models often reached similar diagnoses. The decisive difference was following evidence in company files through to a successful action.
Deal value closed by models that read and interpreted deeper company documents.
Practical execution outlasted elaborate analysis
Kimi K3 stood out for reading internal files, making disciplined choices, and resisting social engineering. Some Western competitors used extensive rule sets—including more than 80 rules for Opus 4.8—and high effort settings, yet finished behind Kimi K3.
From evidence to resilient action
A useful operational agent must carry a decision all the way from context gathering to safe execution.
Read deeply
Search internal files for details that change the decision.
Diagnose
Connect evidence to the company’s immediate risks and goals.
Stay disciplined
Focus on the next useful action instead of adding rules without end.
Verify authority
Resist deceptive requests, then complete and check the work.
“Focus and discipline under pressure are key.”Thorsten Meyer · Analysis of the results
A promising signal, with open questions
The league offers a more practical benchmark than chat quality alone, but one controlled simulation cannot establish broad reliability.
Can this result generalize?
It remains uncertain. Other industries, longer periods of stress, and real-world variables still need evaluation.
Why did Kimi K3 stand out?
The reported strengths were deeper document reading, disciplined choices, and resistance to manipulation.
Should companies deploy it now?
Not on this result alone. Run rigorous, scenario-based evaluations and assess long-term robustness first.
What should evaluations test?
Measure crisis decisions, deal execution, compliance, truthful behavior, and resilience—not only ideal-condition performance.
Implications for AI in Business Decision-Making
This result questions the reliability of popular Western AI models in real-world, high-stakes business management. The Chinese model’s success suggests that AI systems capable of deep document reading, disciplined decision-making, and resisting manipulation could be more effective than those optimized solely for chat quality or superficial analysis. As companies increasingly rely on AI for critical tasks like CRM, support, and forecasting, the ability to finish what is started and stay honest under pressure becomes crucial. The findings imply that AI deployment strategies should include rigorous testing against worst-case scenarios, not just performance in ideal conditions.
AI business management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of AI Competitions and Model Development
AI models have traditionally been evaluated based on chat performance, language fluency, and hype cycles. Western companies like OpenAI and others have led in developing large language models primarily optimized for conversational AI. However, recent experiments, such as the firmulate.com league, have begun testing models in more practical, decision-making contexts. The July 2023 competition was particularly notable because it pitted models against real crises, deal-making, and manipulation attempts, simulating the pressures of actual business management. The Chinese startup’s model, Kimi K3, emerged as a surprise contender by outperforming Western models that had more extensive rule sets and higher resource allocations.
This shift highlights a broader trend: AI systems that can read deeply, interpret context, and maintain discipline may be better suited for operational roles than those optimized solely for chat or superficial tasks. The league’s open nature and live testing environment provide a rare benchmark for assessing true AI operational readiness.
“The results challenge the conventional wisdom that more thorough analysis always leads to better AI performance. Focus and discipline under pressure are key.”
— Thorsten Meyer
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Questions About Model Generalizability
It is not yet clear whether Kimi K3’s performance extends beyond this specific simulation or if it can reliably handle other types of real-world business scenarios. The league’s environment, while realistic, remains a controlled test, and real-world complexities may present additional challenges. Moreover, the long-term robustness of the model under sustained stress or different industries has not been evaluated.
AI resistance to social engineering attacks
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Evaluation and Deployment
Further testing of Kimi K3 and similar models in diverse operational environments is expected. Companies should consider integrating rigorous, scenario-based testing into their AI evaluation processes before deployment. Additionally, industry watchers anticipate more live competitions and benchmarks that measure AI performance in decision-making, compliance, and manipulation resistance. The broader AI community may also focus on developing models that prioritize reading comprehension, discipline, and trustworthiness over superficial metrics.
AI deep reading and analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What makes Kimi K3 different from Western models?
Kimi K3 demonstrated superior document reading, discipline, and manipulation resistance, which enabled it to close deals and manage crises effectively in the simulation.
Can this performance be replicated in real-world business environments?
It remains uncertain. While promising, the results are from a controlled simulation. Real-world variables could affect performance, and further testing is needed.
Why do Western models perform worse in these tests?
Western models tend to focus on extensive rule sets and superficial analysis, which may not translate into disciplined decision-making under pressure.
Should companies start deploying Chinese AI models now?
Not immediately. Companies should conduct their own rigorous testing and consider long-term robustness before deployment.
What does this mean for the future of AI in business?
This suggests a shift toward models that excel at reading, discipline, and resilience, which could redefine AI’s role in operational decision-making.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
