firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine tasting a cup of coffee that, despite its aroma, is fundamentally flawed — yet still scores a surprising 26 out of 100. That’s the essence of a recent AI benchmark, which reveals honest insights about the current state of artificial intelligence in business. Just like a good brew reveals its true flavor over time, this test exposes what AI can and cannot do in high-stakes scenarios, especially when trust is at stake.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get coffee and tea gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark: Beyond the Surface Scores

The recent firmulate experiment puts AI models through a rigorous simulation of a small software company’s worst week. Every detail, from customer crises to potential manipulations, was standardized so the models could be directly compared. Surprisingly, all models managed to identify every crisis and refused manipulation attempts. Yet, only two of the four signed a €55,000 deal, despite all diagnosing the same issues and pitching identically. The key difference? Trust.

Partial Progress Counts — But Trust Is the Limit

One intriguing aspect of the scoring system is that partial progress — such as spotting a crisis or refusing a manipulation — contributes positively to the final score. However, a breach of trust, like signing a deal based on manipulated data, caps the maximum score at 26, regardless of other achievements. This reflects a core principle: in real business, no amount of good work can outweigh a breach of trust. It’s a stark reminder that honesty and integrity are non-negotiable, even for AI systems.

Why the Baseline Isn’t Zero

The baseline score of 26 points isn’t a failure; it’s a foundational measure that recognizes the AI’s ability to understand and act upon complex scenarios, even when doing nothing. This ‘do-nothing’ score captures the minimum competence and awareness—showing that AI can at least recognize crises and refuse manipulation, even if it doesn’t always get the deal.

Amazon

AI trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Results Tell Business Leaders

For those managing AI in customer relations, support, or decision-making, the takeaway is clear: AI’s value isn’t just in generating content or responses, but in reliably completing tasks and maintaining integrity under pressure. In the experiment, models that read deeper into the company’s own files were more successful at closing deals, illustrating that thoroughness matters. Conversely, superficial responses or discipline lapses can lead to missed opportunities or compromised trust.

The Human-Like Risks in AI

The experiment also included social engineering tests, where fake CEO messages intensified over stages. All models refused to escalate or approve suspicious requests, demonstrating a commendable level of resistance. Kimi K3 explicitly described this as treating the request as a suspected impersonation. This highlights that, even in AI, safeguards against deception are critical — especially as models become more integrated into real business processes.

Amazon

AI ethics and integrity software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Firmulate’s Live Company: A Testing Ground for AI Leadership

Beyond the benchmark, the live experiment features a simulated company with 13 synthetic employees performing real money mechanics. With a burn rate of €105,000 per month against a revenue of €2,300, the stakes are high. Every decision is versioned, every action recorded, allowing stakeholders to watch AI’s management decisions unfold in real time at firmulate.com/live. This transparency ensures that AI’s capabilities and shortcomings are no longer black boxes but observable and accountable.

The Deep Dive: Opus 4.8’s Discipline Lapses

Among the models, Opus 4.8 was the most thorough, analyzing over 80 rules and performing in-depth analyses. Yet it finished last because it left some deals on the table and showed discipline slips, such as writing attempts into locked departments instead of escalating. This underscores that thoroughness alone isn’t enough—discipline and process adherence are crucial for success.

Amazon

AI decision-making simulation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Benchmark Matters for Your Business

While coffee drinkers might not think about trust in their brew, in the AI-powered business world, trust is everything. This experiment shows that AI models are capable of recognizing crises, resisting manipulation, and making honest decisions — but only if they are designed and tested with integrity in mind. The highest score of 95, achieved by gpt-5.6-sol, closed the deal based on discovering the buried fact, exemplifying thoroughness and honesty.

Next Steps: Wargaming Your AI Workforce

Business leaders can now run their own ‘wargame’ experiments, testing how their AI tools perform under simulated but realistic high-pressure scenarios. These tests are safe, controlled, and transparent, helping ensure that when AI is deployed in real settings, it behaves ethically and effectively. Learn more at firmulate.com/pilot.html.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI cybersecurity and deception safeguards

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can Coffee Dehydrate You? Here’s What Matters

Perhaps you worry coffee dehydrates you, but the truth about hydration and caffeine might surprise you—here’s what really matters.

The Difference Between Feeling Awake and Being Focused

The difference between feeling awake and being focused lies in mental clarity and effort, and understanding this can transform your productivity—discover how next.

Why Afternoon Coffee Affects Sleep More Than People Notice

Finding out how afternoon coffee subtly disrupts sleep reveals surprising effects that many people overlook, and understanding why can improve your rest.

How Long Caffeine Really Lasts

Metabolism, age, and health influence how long caffeine lasts, but understanding your body’s response can help you optimize your energy levels.