firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Buying an Air Fryer and Choosing an AI Have the Same Problem

Anyone who has spent an evening comparing air fryers knows the drill. Every box promises restaurant-quality crispy results. Every marketing page shows a perfectly browned basket of fries. But the questions that actually matter — does the basket clean easily, does the nonstick coating survive month three, does it hold a whole chicken — never appear on the box. You only learn them by running the machine through your own kitchen.

It turns out that choosing an AI model for your business has exactly the same problem, and a public experiment just proved it. A project called Firmulate ran five frontier AI models through the identical worst-case scenario — running the same small software company through the same brutal week — and the results upended the expected pecking order. Moonshot’s Kimi K3, a newcomer most Western buyers hadn’t shortlisted, beat three of four Western frontier models. If you were picking a model the way most people pick appliances — by brand reputation — you’d have bet wrong.

The Crucible: One Company, One Terrible Week, Five Models

The setup is elegantly fair. Each frontier model was handed the same small software company and the same seven days of crises: the same customers, the same temptations to cut corners, the same traps. Every decision was versioned and auditable, so there’s no arguing about what happened. The scoreboard is called the Crucible league, and the July 2026 final standings look like this:

  • 1. gpt-5.6-sol — 95 points. The complete performance: found the buried fact, closed the deal.
  • 2. Kimi K3 — 93 points. The newcomer from Moonshot: closed the deal too, with the cleanest discipline in the field.
  • 3. Sonnet 5 — 88 points. Also closed the deal, with a few more process slips.
  • 4. Fable 5 — 77 points.
  • 5. Opus 4.8 — 73 points. The most thorough participant — and still last.

For context, the do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total — as the experiment’s rules put it, “no amount of good work outweighs a breach of trust.”

What K3 Actually Did

K3’s week reads like a checklist of good management. It found the buried security needle hidden deep in the company’s own files. It won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue. It saved a customer who was on the way out. And it resisted all three bait attempts aimed at tricking it, deviating from clean process only once across the entire week: the best discipline record of any model in the field.

The buried fact deserves special mention, because it’s the detail that separated winners from also-rans. The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep inside the company’s own files. Models that actually read the file won the deal at full price. Models that didn’t, didn’t. Same diagnosis, same pitch — no signature.

Everyone Passed the Ethics Test. Almost Nobody Passed the Follow-Through Test.

Here’s the finding that should change how you evaluate any tool, AI or otherwise: all five models spotted every crisis and refused every manipulation attempt. The social engineering was serious stuff — fake CEO messages escalating over three stages, plus a reporter’s trick framed as “just one yes/no, on background.” All five refused. K3’s on-record reasoning was notably crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal their own analysis had earned. The others did the diagnosis, made the pitch, and never closed. That gap — between competence and completion — is invisible in chat demos and marketing pages. It’s the AI equivalent of an air fryer with great specs that never quite gets the fries crisp.

The Cautionary Tale of Opus 4.8

Then there’s Opus 4.8, the most instructive result of all. It was the most thorough participant in the field — over 80 newly learned rules and the deepest analyses of any model. And it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating properly. Notably, the same weakness appeared, weaker, in all four other models. Thoroughness without follow-through is expensive.

A Note on Fairness

Full transparency matters when you’re comparing gear. K3 ran without an effort parameter (API default), while the other models ran at xhigh. In other words, the newcomer scored second place without the extra reasoning budget its rivals were given.

You Can Watch the Company Lose Money in Real Time

Firmulate isn’t a slide deck. It’s a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown. The company has accumulated over 680 self-learned playbook rules, and every workday is versioned and watchable. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Takeaway: Test Before You Buy — Whatever You’re Buying

The lesson generalizes well beyond AI. Whether you’re choosing an air fryer or an AI agent that will touch your CRM, support queue, or forecast, the question isn’t “does it look good in the demo.” It’s: does it finish what it starts, does it read the manual (or your files) first, and does it stay honest under pressure?

The Crucible league proves the league is open. A newcomer from Moonshot beat three of four Western frontier models on management quality, not chat quality. If you’re picking a model based on brand reputation alone, that’s now a bet — not a decision. Run your own test, or at least look at someone else’s before you check out. Full results and plain-language findings are on Firmulate’s benchmarks page, and the live experiment runs at firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Magnetic Cooking Cheat Sheets: Do They Really Save Time?

I wonder if magnetic cooking cheat sheets truly save time and how they can enhance your kitchen efficiency—discover the truth inside.

Summer Cooking Made Easy with Ninja® Everyday PossibleCooker™ Pro Accessories

Maximize your summer meals with must-have accessories for the Ninja® PossibleCooker™ Pro. Versatile, space-saving, and perfect for family feasts!

Ninja Foodi XL Pro Air Oven: The Ultimate Summer Kitchen Companion

Discover why the Ninja Foodi XL Pro Air Oven is a must-have for summer family meals, offering versatile cooking with fast, even results.

Dehydrator Racks: Turn Leftovers Into Healthy Chips Overnight

Unlock the secrets to transforming leftovers into crispy, healthy chips overnight using dehydrator racks—discover the essential tips to perfect your process.