
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Discover the Real Test of AI for Business: Not Just Chat, But Results
Imagine an AI that doesn’t just talk a good game, but actually runs a business through its worst week — making decisions, avoiding pitfalls, and closing deals. That’s exactly what the latest experiment by Firmulate reveals, showing a newcomer AI model outperforming several long-established competitors in a real-world simulation.
As an affiliate, we earn on qualifying purchases.
The Live Business Simulation: A True Test of AI Leadership
In July 2026, four frontier AI models faced off in a unique challenge: running a small software company through its worst week. Every decision was real, every crisis genuine, and every temptation to cheat or cut corners was monitored. This was not a simple chat test — it was a comprehensive business simulation that involved real money mechanics, real customers, and the pressure of a live environment.
The League Table: The Results Are In
- gpt-5.6-sol scored the highest at 95, successfully diagnosing buried issues and closing a €55,000 deal.
- Kimi K3, the newcomer from Moonshot, scored just slightly behind at 93, also closing the same deal but with greater discipline and integrity.
- Sonnet 5 followed at 88, managing to close the deal but with some process slips.
- Fable 5 scored 77, and Opus 4.8 lagged at 73, with discipline and thoroughness notably weaker.
Remarkably, all models identified every crisis and refused manipulative social engineering attempts — like fake CEO messages and reporter tricks. The decisive factor? Kimi K3’s ability to read deep into the company’s files and uncover buried information that clinched the deal at full price.
What This Means for Business and AI
This experiment underscores key qualities vital for AI to be truly useful in business: the ability to read and analyze complex data, resist manipulation under pressure, and follow through to completion. It’s no longer just about how convincingly an AI can chat — it’s whether it can deliver real results in tough environments.
The Fairness and Transparency of Testing
It’s important to note that K3 ran without an effort parameter (the API default), while the other models ran at xhigh. This illustrates that even with less aggressive resource settings, the newcomer still performed at an elite level, highlighting its efficiency and robustness.
Watch the Live Experiment
The entire scenario is live and observable at firmulate.com/live. You can see how these AI models make decisions, handle crises, and interact with a simulated real company, making it a valuable resource for managers considering AI automation.
enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and Wellness
While this experiment is business-focused, the underlying message resonates broadly. Just like in health and wellness, where the goal is to achieve real, measurable improvements, in AI-driven business processes, the proof is in the results — not just in the quality of conversation. A tool that can consistently finish what it starts, analyze deeply, and resist shortcuts is worth more than one that merely impresses in demos.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
