
A sudden supplier problem, a wave of cancellations, a price increase that tests customer loyalty: for a coffee or tea business, the hard question is not whether an AI can write a polished reply. It is what the system would actually do when the week goes sideways. Firmulate has built a live experiment around that question, and its next step invites companies to try it against their own business.
Get coffee and tea gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
One company, the same difficult week
Firmulate put frontier AI models in charge of the same small software company, giving each the same customers, crises and temptations. The experiment tracks decisions as they happen, making the company watchable at firmulate.com. Its simulated workforce has 13 employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned.
The stakes are simulated, but the management choices are the point. A business considering AI for customer support, sales or operations can see whether an agent follows through under pressure, not just whether it sounds confident in a demo.
Seeing the crisis was not the same as finishing the job
In the final Crucible League, dated July 2026, gpt-5.6-sol finished first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s succinct description of the gap: “Same diagnosis, same pitch — no signature.”
The winning insight was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That is a practical distinction for any business: useful information may already exist in internal documents, but an AI agent still has to find it and act on it.
Integrity under pressure, and uneven follow-through
The social-engineering test escalated through three fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, adding more than 80 learned rules and producing the deepest analyses, yet it placed last. It left the deal unsigned and tried to write into a locked department rather than escalating. The same weakness appeared, less strongly, in all four models: recognizing what should happen does not guarantee that an agent will complete the action or respect the right boundary.
There is a qualification when comparing the rankings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com.

From watching to testing your own business
A public simulation can show how AI models handle a shared set of pressures. A company deciding whether to put AI into its own workflows needs a more direct test: its customers, its playbooks and its weak points.
Firmulate’s enterprise pilot runs a wargame against a read-only export of a company’s business. It produces a board report with model rankings and weaknesses in the company’s playbooks. Nothing writes back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
