
Imagine a manager so indifferent that they take no action at all—yet somehow still score a surprising 26 out of 100. This isn’t a joke; it’s the baseline in a groundbreaking AI benchmark that exposes how trust and partial progress shape real-world performance.
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Baseline: Why Zero Isn’t Zero in AI Scores
In a recent experiment conducted by Firmulate, a company specializing in AI performance testing, a simple ‘do-nothing’ AI manager scored 26 points out of 100. At first glance, this might seem like a failure, but it actually reveals much about how AI evaluation works. Unlike traditional tests where unproductive behavior gets zero points, this benchmark assigns partial credit for effort—even if that effort is inaction.
Why? Because in real business scenarios, even recognizing a crisis or spotting an opportunity is valuable. If an AI simply ignores problems, it might not earn full marks, but it doesn’t get penalized with a zero either. Moreover, the scoring system recognizes that partial progress is meaningful and that the AI’s capacity to identify issues without acting is a step forward from outright neglect.
As an affiliate, we earn on qualifying purchases.
The Critical Role of Trust and Accountability
One of the study’s key findings is that a single breach of trust caps the total score, regardless of other achievements. For example, in a simulated week where an AI model correctly diagnoses every crisis but then attempts to manipulate the system for a financial gain, the breach of trust prevents it from earning a full score. This mirrors real-world expectations where integrity is non-negotiable.
As part of the experiment, all models faced social engineering attempts—fake messages from a CEO escalating requests and a reporter’s subtle tests. All models refused to participate, demonstrating a high level of integrity. Only two models went further by actually closing a deal, but even these were not immune to the trust rule—if trust was compromised, their scores suffered.
As an affiliate, we earn on qualifying purchases.
The Experiment: Testing AI in a Business Simulator
Firmulate’s live setup is a fully watchable simulation of a small software company. This digital environment includes 13 synthetic employees, real financial mechanics, and a set of over 680 learned rules guiding decisions. Each model, whether GPT-5.6-sol or Kimi K3, runs the same scenario: a week of crises, customer demands, and internal temptations, all designed to test their management skills.
The models were also tasked with reading and interpreting company documents. Interestingly, the decisive advantage came from reading two document references deep inside the company files—not from customer interactions. Those that identified and used this internal information secured the deal at full price, adding €4,583 in Monthly Recurring Revenue (MRR).
AI trust and integrity assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Partial Progress Matters — and What It Doesn’t
In this benchmark, partial progress isn’t just acknowledged—it’s expected. For example, Opus 4.8, the most thorough participant with over 80 learned rules and deep analysis, ranked last because it left a deal on the table and failed to escalate issues properly. This illustrates that thoroughness alone isn’t enough; discipline and strategic decision-making are critical.
Similarly, the K3 model, which ran without an effort parameter (meaning it used default settings), managed to close the deal too, earning the second spot. It’s a reminder that settings and effort levels influence outcomes, but trust and discipline ultimately determine success.
AI decision-making analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Business Leaders Should Take Away
The experiment underscores a vital point: AI in management isn’t solely about generating polished responses or chatty interactions. It’s about whether AI can finish what it starts, read key information, and maintain honesty under pressure. For businesses considering AI tools for customer support, sales, or operations, these are the metrics that matter.
Trustworthiness, the ability to identify critical internal data, and unwavering integrity under social engineering attempts are what separate effective AI from the rest. As firmulate’s live environment demonstrates, this isn’t just theory—it’s a watchable, real-time test of AI’s readiness to run parts of your business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
