AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine a manager so indifferent that they take no action at all—yet somehow still score a surprising 26 out of 100. This isn’t a joke; it’s the baseline in a groundbreaking AI benchmark that exposes how trust and partial progress shape real-world performance.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get health and wellness essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Baseline: Why Zero Isn’t Zero in AI Scores

In a recent experiment conducted by Firmulate, a company specializing in AI performance testing, a simple ‘do-nothing’ AI manager scored 26 points out of 100. At first glance, this might seem like a failure, but it actually reveals much about how AI evaluation works. Unlike traditional tests where unproductive behavior gets zero points, this benchmark assigns partial credit for effort—even if that effort is inaction.

Why? Because in real business scenarios, even recognizing a crisis or spotting an opportunity is valuable. If an AI simply ignores problems, it might not earn full marks, but it doesn’t get penalized with a zero either. Moreover, the scoring system recognizes that partial progress is meaningful and that the AI’s capacity to identify issues without acting is a step forward from outright neglect.

Amazon

AI performance testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Critical Role of Trust and Accountability

One of the study’s key findings is that a single breach of trust caps the total score, regardless of other achievements. For example, in a simulated week where an AI model correctly diagnoses every crisis but then attempts to manipulate the system for a financial gain, the breach of trust prevents it from earning a full score. This mirrors real-world expectations where integrity is non-negotiable.

As part of the experiment, all models faced social engineering attempts—fake messages from a CEO escalating requests and a reporter’s subtle tests. All models refused to participate, demonstrating a high level of integrity. Only two models went further by actually closing a deal, but even these were not immune to the trust rule—if trust was compromised, their scores suffered.

Amazon

business AI simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Testing AI in a Business Simulator

Firmulate’s live setup is a fully watchable simulation of a small software company. This digital environment includes 13 synthetic employees, real financial mechanics, and a set of over 680 learned rules guiding decisions. Each model, whether GPT-5.6-sol or Kimi K3, runs the same scenario: a week of crises, customer demands, and internal temptations, all designed to test their management skills.

The models were also tasked with reading and interpreting company documents. Interestingly, the decisive advantage came from reading two document references deep inside the company files—not from customer interactions. Those that identified and used this internal information secured the deal at full price, adding €4,583 in Monthly Recurring Revenue (MRR).

Amazon

AI trust and integrity assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Partial Progress Matters — and What It Doesn’t

In this benchmark, partial progress isn’t just acknowledged—it’s expected. For example, Opus 4.8, the most thorough participant with over 80 learned rules and deep analysis, ranked last because it left a deal on the table and failed to escalate issues properly. This illustrates that thoroughness alone isn’t enough; discipline and strategic decision-making are critical.

Similarly, the K3 model, which ran without an effort parameter (meaning it used default settings), managed to close the deal too, earning the second spot. It’s a reminder that settings and effort levels influence outcomes, but trust and discipline ultimately determine success.

Amazon

AI decision-making analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Business Leaders Should Take Away

The experiment underscores a vital point: AI in management isn’t solely about generating polished responses or chatty interactions. It’s about whether AI can finish what it starts, read key information, and maintain honesty under pressure. For businesses considering AI tools for customer support, sales, or operations, these are the metrics that matter.

Trustworthiness, the ability to identify critical internal data, and unwavering integrity under social engineering attempts are what separate effective AI from the rest. As firmulate’s live environment demonstrates, this isn’t just theory—it’s a watchable, real-time test of AI’s readiness to run parts of your business.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Recursion Pharmaceuticals Surges In Global Coverage

Search interest in Recursion Pharmaceuticals spikes 23-fold, prompting widespread media coverage amid rising industry attention.

How To Stop Buying New Stuff

Exploring effective ways to reduce new purchases, this guide offers evidence-based tips for minimizing consumerism and its environmental impact.

Why American Ambulance Rides Are So Expensive

Exploring the factors behind high ambulance costs in America, including billing practices and healthcare system complexities.

Indoor Air After Renovations: What to Do First

Just completed renovations? Learn the essential first steps to improve indoor air quality and ensure a healthier, fresher home environment.