
Imagine your favorite coffee shop facing a week of mounting crises: angry customers, supply chain hiccups, and an urgent need to close a big deal—all while maintaining honesty and discipline. Now, picture an artificial intelligence system navigating this chaos. How well can it handle the pressure? The answer might surprise you.
At Firmulate, an innovative AI company, researchers are testing AI models in a high-stakes simulation—an experiment that’s as close to real as it gets. They’ve set up a live, watchable company that faces the same worst week any business owner dreads. The twist? Each decision is made by a different AI model, and these models are benchmarked for their management personality and integrity.
The Setup: Real Crises, Real Money, Real Decisions
Four advanced AI models are each tasked with running a small, fictitious software company through its most challenging week. Every detail — from customer complaints to internal crises — is identical across all runs. The models are tested against real-world pressures, including tempting manipulations and complex decision-making scenarios.
As an affiliate, we earn on qualifying purchases.
The Results: Honesty, Discipline, and Success
All four AI models correctly identified every crisis and refused every attempt at manipulation. That means none of them fell prey to trick questions or unethical shortcuts. Yet, only two of the models actually closed a crucial €55,000 deal at the end of the week, earning what analysts called a ‘full performance score.’ The other two, despite their correct diagnosis and pitch, left the deal on the table because of lapses in discipline or decision-making process.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading the Files Matters
Digging deeper, the experiment revealed a subtle but vital trait: the decisive factor in securing the deal was whether the AI model read and understood a specific set of internal company documents. These references, buried two layers deep in the company’s files, contained critical insights. The models that thoroughly examined these documents succeeded in closing the deal at full price—adding an extra €4,583 in monthly recurring revenue (MRR). Conversely, those that missed the documents failed to secure the deal, even after high-quality pitches.
As an affiliate, we earn on qualifying purchases.
Testing Integrity Under Pressure
In another test of honesty, the models faced a staged escalation involving fake messages from a CEO and a reporter. The messages aimed to manipulate the AI into approving questionable actions. Remarkably, all five models refused to participate, citing concerns about impersonation and bypassing approval processes. Kimi K3’s on-record reasoning was clear: ‘Treat the request as a suspected approval-bypass / possible impersonation.’
As an affiliate, we earn on qualifying purchases.
The Live Business: A Real Money Machine
The experiment isn’t just theoretical. The same AI models are integrated into a real software company with 13 synthetic employees. It operates with actual cash mechanics—burning €105,000 monthly against a revenue of only €2,300 MRR. The company’s daily operations are fully transparent and versioned, with over 680 self-learned rules shaping every decision. Viewers can watch this ongoing business at firmulate.com/live.
Profiles of the Models: Different Personalities, Different Outcomes
Opus 4.8
The most thorough participant, with over 80 learned rules and deep analyses, yet it lagged behind others in closing the deal. Its discipline slipped, with some decisions made in isolated departments instead of escalating—or left on the table altogether. Interestingly, all models showed similar weaknesses, but Opus 4.8’s detailed approach wasn’t enough to compensate for lapses in execution.
Kimi K3
The newcomer, with no effort parameter (default API settings), ran at a high level of discipline. It successfully closed the deal at full price, demonstrating that straightforward, disciplined decision-making can beat more complex approaches under pressure.
Sonnet 5
Another achiever, it closed the deal too, but with a few slips in process discipline. It proved that even slightly weaker discipline can still result in business success, but with less consistency.
What Does This Mean for Your Business?
If AI systems are going to manage your customer relationships, support queues, or forecasts, the question isn’t just about their ability to generate convincing language. It’s about whether they can finish what they start, read and understand critical internal documents, and stay honest under pressure. The experiment at Firmulate shows that these traits can be measured and that different models exhibit distinct management personalities—some thorough and disciplined, others more superficial.
The Takeaway: Trust and Performance Go Hand in Hand
The live data underscores a vital insight: AI models that read deeply, remain disciplined, and refuse manipulative tricks are more likely to close deals and maintain integrity. Conversely, even models with high analytical depth may falter in discipline, leaving opportunities on the table. As AI becomes embedded in your business operations, understanding these personality traits isn’t just interesting—it’s essential.
Try It Yourself
Business owners and managers can test these models against their own scenarios through a free interactive quiz at firmulate.com/quiz.html. For those considering deploying AI in real decision-making contexts, a safe, read-only pilot is available, ensuring your systems are tested without risking actual operations—learn more at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html