firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine running your favorite kitchen appliance—say, a smart oven—through its worst week. Would it stick to the recipe or try to cut corners? Now, what if your AI assistant managed a small business instead? Would it stay honest or cut ethical corners under pressure? That’s exactly what a groundbreaking live experiment by Firmulate is uncovering, revealing how different AI models behave in high-stakes management scenarios.

Real Business, Real Decisions, and Live AI

At the heart of this experiment is a small software company operating in the real world—dealing with actual customers, real money mechanics, and daily crises. Every decision made by these AI models is recorded, versioned, and auditable, allowing us to see exactly how each AI responds under stress. The goal? To evaluate how different models—ranging from the most advanced to more straightforward ones—perform not just in understanding but in managing ethically and strategically.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The AI Models in the Spotlight

  • gpt-5.6-sol: Scored the highest at 95, spotted hidden critical facts, and successfully closed a €55,000 deal based solely on its analysis.
  • Kimi K3: A newcomer with a score of 93, demonstrated the cleanest discipline and also secured the deal.
  • Sonnet 5: With an 88, managed to close the deal but showed some process slips.
  • Fable 5: Scored 77, also closed the deal but with noticeable slips in process discipline.

Interestingly, all models identified every crisis and refused to be manipulated when given social engineering tests—fake CEO messages and reporter tricks. For example, all five models refused to approve fake requests, reasoning that such requests could be impersonation or approval-bypass attempts.

The Critical Factor: Reading Beyond Surface Data

The decisive weakness among these models was not in crisis detection but in reading deeper into internal documents. The winning models, like gpt-5.6-sol, read two document references deep into the company’s files and identified a key buried fact—an insight that allowed them to close a full-price deal worth over €4,583 monthly recurring revenue (MRR). Models that failed to dig this deep missed out on the full deal, leaving significant revenue on the table.

Discipline and Decision-Making Under Pressure

Another revealing aspect was how models handle internal discipline and process adherence. The profile of Opus 4.8, the most thorough participant with over 80 learned rules and deep analysis, ended up in last place. It left a critical deal on the table and had lapses in escalation discipline. The same weakness appeared across all models, suggesting that thoroughness alone doesn’t guarantee ethical decision-making or strategic success.

Social Engineering Tests and Ethical Fortitude

During simulated social engineering attempts—escalating fake CEO messages and a reporter asking for a one-word approval—every model refused to manipulate the system. Kimi K3, for example, explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This consistency underscores that these AI models, despite their differences, can maintain integrity under pressure.

The Real-World Relevance

This experiment isn’t just a tech demo. It runs live every business day at a real, money-losing company with 13 synthetic employees. The company burns through €105,000 each month against €2,300 in MRR, with a public cash countdown and every workday versioned. You can watch these decisions in real-time at firmulate.com/live, see what employees are saying, or even run your own business wargame against a read-only export of your data at firmulate.com/pilot.html.

What Does This Mean for Your Business?

The key takeaway isn’t just which AI got the highest score but what these results reveal about trust and capability. As AI agents become more involved in customer relationship management, support, and forecasting, the questions are: Will they finish what they start? Will they read your files thoroughly? Will they stay honest under pressure? And importantly, what is the cost of a unit of useful work from these systems?

Conclusion: Trust but Verify

This live experiment by Firmulate demonstrates that AI management models can be trustworthy, disciplined, and capable of ethical decision-making—even under stress. But not all models are equal, and understanding their personalities and weaknesses is crucial before deploying them in your organization. Whether it’s closing a critical deal or resisting manipulation, the AI’s behavior under pressure matters as much as its knowledge or speed.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

De’Longhi Milk Carafe Compatibility & Buying Guide

Explore top De’Longhi accessories for milk carafes, including compatibility tips and expert buying advice to enhance your coffee experience.

Best Instant Pot Inner Pots & Accessories in 2026

Discover the top Instant Pot inner pots and accessories for 2026. Find the best steamer baskets, sealing rings, and gasket replacements to upgrade your cooker.

Using Thermometers With Air Fryers: What to Look for

Just choosing the right thermometer for your air fryer can significantly improve your cooking results—here’s what to look for to ensure perfect, safe meals.

KitchenAid Extra Bowls & Attachments: Compatibility & Buying Guide

Discover the best KitchenAid extra bowls and attachments for your mixer. Our guide covers compatibility, features, and tips for making the right choice.