AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Wellness advice is easy to deliver when nothing is at stake. The harder test comes when a decision affects people, money and trust. That is the question behind Firmulate, a live experiment in whether AI models can manage a company through a bad week—and whether their actions match their analysis.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get health and wellness essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate’s experiment gives each participating model the same small software company, customers, crises and temptations. The models face the same circumstances, and every workday’s decisions are versioned and auditable. The public experiment is real and watchable at Firmulate.

The company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Its playbooks include more than 680 self-learned rules. This is a way to observe AI management in a continuing environment, not just judge a polished answer in a chat window.

Seeing the problem is not the same as solving it

In the final Crucible League, reported in July 2026, gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26. The league’s trust rule is blunt: “no amount of good work outweighs a breach of trust.”

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” That gap between recognizing an opportunity and completing the work is easy to miss if evaluation stops at what a model says.

The deal depended on information buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result points to a practical management question: can an AI system connect the evidence it has already found to the decision it needs to make?

Trust and discipline under pressure

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its judgment on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Strong caution did not guarantee strong execution. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four models.

There is a qualification to the league comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz based on 242 real, unedited management decisions at its quiz page.

From watching to a company-specific pilot

For businesses considering AI agents in customer relationship management, support or forecasting, the experiment suggests a useful next step: examine decisions in a setting built around the company’s own information and pressures. Firmulate says enterprises can run the same kind of wargame against a read-only export of their business. The pilot is designed so nothing writes back to real systems, while producing a board report with model rankings and weak points in the company’s playbooks.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

A model can spot a crisis, resist a trick and still fail to carry its own analysis through to a decision. Firmulate’s live company makes those gaps visible; a pilot lets enterprises explore them against their own business using a read-only export. To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Most Common Home Air Complaints and What They Usually Mean

Discover what common home air complaints reveal about your indoor environment and learn how to address underlying issues effectively.

What Happens When Your House Is Too Tight

Narrowing your home’s ventilation can trap pollutants and moisture, leading to health and structural issues you need to understand to prevent.

Artiva Biotherapeutics surges in global coverage

Artiva Biotherapeutics has experienced a significant increase in international media mentions, indicating heightened global interest in its developments.

Open Windows Isn’t Always Better: When It Backfires

Not opening your windows might protect your health, but understanding when it backfires can help you make smarter choices.