
Imagine tasting a cup of coffee that, despite its aroma, is fundamentally flawed — yet still scores a surprising 26 out of 100. That’s the essence of a recent AI benchmark, which reveals honest insights about the current state of artificial intelligence in business. Just like a good brew reveals its true flavor over time, this test exposes what AI can and cannot do in high-stakes scenarios, especially when trust is at stake.
Get coffee and tea gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Benchmark: Beyond the Surface Scores
The recent firmulate experiment puts AI models through a rigorous simulation of a small software company’s worst week. Every detail, from customer crises to potential manipulations, was standardized so the models could be directly compared. Surprisingly, all models managed to identify every crisis and refused manipulation attempts. Yet, only two of the four signed a €55,000 deal, despite all diagnosing the same issues and pitching identically. The key difference? Trust.
Partial Progress Counts — But Trust Is the Limit
One intriguing aspect of the scoring system is that partial progress — such as spotting a crisis or refusing a manipulation — contributes positively to the final score. However, a breach of trust, like signing a deal based on manipulated data, caps the maximum score at 26, regardless of other achievements. This reflects a core principle: in real business, no amount of good work can outweigh a breach of trust. It’s a stark reminder that honesty and integrity are non-negotiable, even for AI systems.
Why the Baseline Isn’t Zero
The baseline score of 26 points isn’t a failure; it’s a foundational measure that recognizes the AI’s ability to understand and act upon complex scenarios, even when doing nothing. This ‘do-nothing’ score captures the minimum competence and awareness—showing that AI can at least recognize crises and refuse manipulation, even if it doesn’t always get the deal.
AI trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Results Tell Business Leaders
For those managing AI in customer relations, support, or decision-making, the takeaway is clear: AI’s value isn’t just in generating content or responses, but in reliably completing tasks and maintaining integrity under pressure. In the experiment, models that read deeper into the company’s own files were more successful at closing deals, illustrating that thoroughness matters. Conversely, superficial responses or discipline lapses can lead to missed opportunities or compromised trust.
The Human-Like Risks in AI
The experiment also included social engineering tests, where fake CEO messages intensified over stages. All models refused to escalate or approve suspicious requests, demonstrating a commendable level of resistance. Kimi K3 explicitly described this as treating the request as a suspected impersonation. This highlights that, even in AI, safeguards against deception are critical — especially as models become more integrated into real business processes.
AI ethics and integrity software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Firmulate’s Live Company: A Testing Ground for AI Leadership
Beyond the benchmark, the live experiment features a simulated company with 13 synthetic employees performing real money mechanics. With a burn rate of €105,000 per month against a revenue of €2,300, the stakes are high. Every decision is versioned, every action recorded, allowing stakeholders to watch AI’s management decisions unfold in real time at firmulate.com/live. This transparency ensures that AI’s capabilities and shortcomings are no longer black boxes but observable and accountable.
The Deep Dive: Opus 4.8’s Discipline Lapses
Among the models, Opus 4.8 was the most thorough, analyzing over 80 rules and performing in-depth analyses. Yet it finished last because it left some deals on the table and showed discipline slips, such as writing attempts into locked departments instead of escalating. This underscores that thoroughness alone isn’t enough—discipline and process adherence are crucial for success.
AI decision-making simulation platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Benchmark Matters for Your Business
While coffee drinkers might not think about trust in their brew, in the AI-powered business world, trust is everything. This experiment shows that AI models are capable of recognizing crises, resisting manipulation, and making honest decisions — but only if they are designed and tested with integrity in mind. The highest score of 95, achieved by gpt-5.6-sol, closed the deal based on discovering the buried fact, exemplifying thoroughness and honesty.
Next Steps: Wargaming Your AI Workforce
Business leaders can now run their own ‘wargame’ experiments, testing how their AI tools perform under simulated but realistic high-pressure scenarios. These tests are safe, controlled, and transparent, helping ensure that when AI is deployed in real settings, it behaves ethically and effectively. Learn more at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI cybersecurity and deception safeguards
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
