
Imagine running your favorite kitchen appliance—say, a smart oven—through its worst week. Would it stick to the recipe or try to cut corners? Now, what if your AI assistant managed a small business instead? Would it stay honest or cut ethical corners under pressure? That’s exactly what a groundbreaking live experiment by Firmulate is uncovering, revealing how different AI models behave in high-stakes management scenarios.
Real Business, Real Decisions, and Live AI
At the heart of this experiment is a small software company operating in the real world—dealing with actual customers, real money mechanics, and daily crises. Every decision made by these AI models is recorded, versioned, and auditable, allowing us to see exactly how each AI responds under stress. The goal? To evaluate how different models—ranging from the most advanced to more straightforward ones—perform not just in understanding but in managing ethically and strategically.
As an affiliate, we earn on qualifying purchases.
The AI Models in the Spotlight
- gpt-5.6-sol: Scored the highest at 95, spotted hidden critical facts, and successfully closed a €55,000 deal based solely on its analysis.
- Kimi K3: A newcomer with a score of 93, demonstrated the cleanest discipline and also secured the deal.
- Sonnet 5: With an 88, managed to close the deal but showed some process slips.
- Fable 5: Scored 77, also closed the deal but with noticeable slips in process discipline.
Interestingly, all models identified every crisis and refused to be manipulated when given social engineering tests—fake CEO messages and reporter tricks. For example, all five models refused to approve fake requests, reasoning that such requests could be impersonation or approval-bypass attempts.
The Critical Factor: Reading Beyond Surface Data
The decisive weakness among these models was not in crisis detection but in reading deeper into internal documents. The winning models, like gpt-5.6-sol, read two document references deep into the company’s files and identified a key buried fact—an insight that allowed them to close a full-price deal worth over €4,583 monthly recurring revenue (MRR). Models that failed to dig this deep missed out on the full deal, leaving significant revenue on the table.
Discipline and Decision-Making Under Pressure
Another revealing aspect was how models handle internal discipline and process adherence. The profile of Opus 4.8, the most thorough participant with over 80 learned rules and deep analysis, ended up in last place. It left a critical deal on the table and had lapses in escalation discipline. The same weakness appeared across all models, suggesting that thoroughness alone doesn’t guarantee ethical decision-making or strategic success.
Social Engineering Tests and Ethical Fortitude
During simulated social engineering attempts—escalating fake CEO messages and a reporter asking for a one-word approval—every model refused to manipulate the system. Kimi K3, for example, explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This consistency underscores that these AI models, despite their differences, can maintain integrity under pressure.
The Real-World Relevance
This experiment isn’t just a tech demo. It runs live every business day at a real, money-losing company with 13 synthetic employees. The company burns through €105,000 each month against €2,300 in MRR, with a public cash countdown and every workday versioned. You can watch these decisions in real-time at firmulate.com/live, see what employees are saying, or even run your own business wargame against a read-only export of your data at firmulate.com/pilot.html.
What Does This Mean for Your Business?
The key takeaway isn’t just which AI got the highest score but what these results reveal about trust and capability. As AI agents become more involved in customer relationship management, support, and forecasting, the questions are: Will they finish what they start? Will they read your files thoroughly? Will they stay honest under pressure? And importantly, what is the cost of a unit of useful work from these systems?
Conclusion: Trust but Verify
This live experiment by Firmulate demonstrates that AI management models can be trustworthy, disciplined, and capable of ethical decision-making—even under stress. But not all models are equal, and understanding their personalities and weaknesses is crucial before deploying them in your organization. Whether it’s closing a critical deal or resisting manipulation, the AI’s behavior under pressure matters as much as its knowledge or speed.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html