firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A sudden supplier problem, a wave of cancellations, a price increase that tests customer loyalty: for a coffee or tea business, the hard question is not whether an AI can write a polished reply. It is what the system would actually do when the week goes sideways. Firmulate has built a live experiment around that question, and its next step invites companies to try it against their own business.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get coffee and tea gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

One company, the same difficult week

Firmulate put frontier AI models in charge of the same small software company, giving each the same customers, crises and temptations. The experiment tracks decisions as they happen, making the company watchable at firmulate.com. Its simulated workforce has 13 employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

The stakes are simulated, but the management choices are the point. A business considering AI for customer support, sales or operations can see whether an agent follows through under pressure, not just whether it sounds confident in a demo.

Seeing the crisis was not the same as finishing the job

In the final Crucible League, dated July 2026, gpt-5.6-sol finished first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s succinct description of the gap: “Same diagnosis, same pitch — no signature.”

The winning insight was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That is a practical distinction for any business: useful information may already exist in internal documents, but an AI agent still has to find it and act on it.

Integrity under pressure, and uneven follow-through

The social-engineering test escalated through three fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.”

Opus 4.8 was the most thorough participant, adding more than 80 learned rules and producing the deepest analyses, yet it placed last. It left the deal unsigned and tried to write into a locked department rather than escalating. The same weakness appeared, less strongly, in all four models: recognizing what should happen does not guarantee that an agent will complete the action or respect the right boundary.

There is a qualification when comparing the rankings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to testing your own business

A public simulation can show how AI models handle a shared set of pressures. A company deciding whether to put AI into its own workflows needs a more direct test: its customers, its playbooks and its weak points.

Firmulate’s enterprise pilot runs a wargame against a read-only export of a company’s business. It produces a board report with model rankings and weaknesses in the company’s playbooks. Nothing writes back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Sensory Basics Behind Tasting Coffee Better

Irritated by dull coffee flavors? Unlock the sensory secrets that will elevate your tasting experience and deepen your appreciation for each sip.

The AI Model Taste Test Has a Surprise Second Place

Moonshot’s Kimi K3 placed second in Firmulate’s company-management trial, ahead of three rivals. The results suggest businesses should test models on their own work.

AI Models Proven Resilient in Simulated Corporate Crisis — No Compromises on Integrity

Leading AI models have demonstrated their ability to resist manipulation and uphold integrity in simulated crises, proving trustworthiness before real-world deployment.

Does Dark Roast Really Have Less Caffeine?

AIThis post was created with the assistance of artificial intelligence (AI).No, dark…