
Get coffee and tea gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A blind taste test for AI managers
Coffee drinkers know that a familiar label doesn’t guarantee a better cup. A new workplace experiment makes a similar case about AI: when models were asked to run the same small software company through its worst week, a newcomer from Moonshot finished ahead of three established Western rivals. The result is a reminder that choosing an AI model by reputation alone can be a costly bet.
From chat to company management
Firmulate’s Crucible put frontier AI models in charge of the same customers, crises and temptations. The company is software, but the stakes in the exercise are designed to feel like work: decisions are versioned and auditable, and the models have to handle a week’s worth of competing demands rather than simply produce a polished answer.
In the final league table for July 2026, gpt-5.6-sol took first place with 95. Moonshot’s Kimi K3 came second with 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The experiment’s rule is stark: partial progress counts, but one breach of trust caps the total. As Firmulate puts it, “no amount of good work outweighs a breach of trust.”
The headline result is close, but the work behind it matters more than the leaderboard alone. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The company’s files contained a competitor weakness buried two document references deep. Models that found it won the deal at full price, worth +€4,583 MRR. Others could diagnose the opportunity and make the pitch, then leave the signature unfinished: “Same diagnosis, same pitch — no signature.”
Pressure, judgment and follow-through
The test also included fake CEO messages escalating across three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a useful distinction for businesses weighing automation: sounding confident is one thing; resisting a plausible shortcut while protecting the company is another.
Opus 4.8 illustrates why a strong impression can be misleading. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of the same discipline problem appeared in all four. The exercise therefore asks more than whether an AI can reason through a crisis. It asks whether it can carry the reasoning through to a sound, authorized action.
The company itself is presented as a live experiment, not a slide deck. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and versioned workdays. Readers can watch the company live or see the benchmark findings. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.
There is an important fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh. That difference belongs alongside the result when interpreting the standings. Firmulate also says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems.

Test before you choose
K3’s second-place finish makes the field look open, but the more practical lesson is that capability depends on the task. In this trial, finding a fact buried in company documents and completing a high-value decision separated models that otherwise recognized the same crises and rejected the same baits. For businesses considering AI in customer support, CRM or forecasting, a model’s brand is only a starting point. The useful question is how it behaves in your own company’s difficult week.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
