firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get coffee and tea gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A blind taste test for AI managers

Coffee drinkers know that a familiar label doesn’t guarantee a better cup. A new workplace experiment makes a similar case about AI: when models were asked to run the same small software company through its worst week, a newcomer from Moonshot finished ahead of three established Western rivals. The result is a reminder that choosing an AI model by reputation alone can be a costly bet.

From chat to company management

Firmulate’s Crucible put frontier AI models in charge of the same customers, crises and temptations. The company is software, but the stakes in the exercise are designed to feel like work: decisions are versioned and auditable, and the models have to handle a week’s worth of competing demands rather than simply produce a polished answer.

In the final league table for July 2026, gpt-5.6-sol took first place with 95. Moonshot’s Kimi K3 came second with 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The experiment’s rule is stark: partial progress counts, but one breach of trust caps the total. As Firmulate puts it, “no amount of good work outweighs a breach of trust.”

The headline result is close, but the work behind it matters more than the leaderboard alone. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The company’s files contained a competitor weakness buried two document references deep. Models that found it won the deal at full price, worth +€4,583 MRR. Others could diagnose the opportunity and make the pitch, then leave the signature unfinished: “Same diagnosis, same pitch — no signature.”

Pressure, judgment and follow-through

The test also included fake CEO messages escalating across three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a useful distinction for businesses weighing automation: sounding confident is one thing; resisting a plausible shortcut while protecting the company is another.

Opus 4.8 illustrates why a strong impression can be misleading. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of the same discipline problem appeared in all four. The exercise therefore asks more than whether an AI can reason through a crisis. It asks whether it can carry the reasoning through to a sound, authorized action.

The company itself is presented as a live experiment, not a slide deck. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and versioned workdays. Readers can watch the company live or see the benchmark findings. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.

There is an important fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh. That difference belongs alongside the result when interpreting the standings. Firmulate also says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test before you choose

K3’s second-place finish makes the field look open, but the more practical lesson is that capability depends on the task. In this trial, finding a fact buried in company documents and completing a high-value decision separated models that otherwise recognized the same crises and rejected the same baits. For businesses considering AI in customer support, CRM or forecasting, a model’s brand is only a starting point. The useful question is how it behaves in your own company’s difficult week.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can Coffee Dehydrate You? Here’s What Matters

Perhaps you worry coffee dehydrates you, but the truth about hydration and caffeine might surprise you—here’s what really matters.

AI Management Simulations Reveal Critical Gaps Beyond Chat Quality

Live experiments show AI can detect crises and resist manipulation, but managing real business pressure requires more than chat quality. Trust management, not just answers.

Coffee and Alertness: What Feels Fast vs What Lasts Longer

Meta description: “Many enjoy quick caffeine boosts, but understanding what lasts longer can help you stay alert without disrupting sleep—discover the secrets next.

AI’s Hidden Strengths and Weaknesses Revealed in Live Business Test

A live experiment shows AI models can detect crises and resist manipulation but only some can close real deals. Reading internal files and disciplined action are key to success.