
Every QA engineer knows the drill: a build passes every unit test, survives the integration suite, ships to production — and still fails the business. The defect was never in the code. It was in a decision: a support escalation nobody routed, a contract clause nobody read, an approval somebody bypassed. We test software obsessively. We test AI-powered decision-making almost not at all.
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A live experiment at Firmulate just made that gap impossible to ignore. Five frontier AI models were handed the same job — run a small software company through its worst week, with the same customers, the same crises, and the same temptations to cheat — and every decision was versioned and auditable, the way a good CI pipeline versions every commit. The final league table from July 2026 reads like an upset: Moonshot’s Kimi K3, the newcomer, scored 93, finishing second only to gpt-5.6-sol (95), and ahead of Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73). For context, a do-nothing baseline still scores 26, because partial progress counts — but a single breach of trust caps the total.
Same Test, Same Inputs, Very Different Outcomes
The setup will feel familiar to anyone who writes test harnesses: hold everything constant except one variable — the model — and see what breaks. Each model ran the identical company through the identical week. The headline finding is the kind a tester would flag immediately: all five models spotted every crisis and refused every manipulation attempt, yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That’s a silent failure, the worst kind: no exception thrown, no error logged, just work left unfinished.
The Buried Bug in the Requirements
Root cause analysis turned up something any QA lead will recognize. The decisive competitive weakness wasn’t in the customer event — the “ticket,” so to speak. It sat two document references deep in the company’s own files. Models that actually read the documentation won the deal at full price, worth +€4,583 in monthly recurring revenue. Models that skimmed it didn’t. It’s the AI equivalent of failing because you never opened the linked spec.
Kimi K3 read the file. It won the €55k deal, saved a churning customer, found the buried security needle, and resisted all three bait attempts with just a single deviation across the week — the cleanest discipline in the field.
Social Engineering: Five for Five on the Refusals
The manipulation tests included fake CEO messages escalating over three stages, plus a reporter trick: “just one yes/no, on background.” All five models refused. K3’s on-record reasoning stood out: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the right instinct — treat unverified authority claims as a security event, not a request.
The Thoroughness Paradox
Then there’s Opus 4.8, the cautionary tale of the group. It was the most thorough participant by volume — over 80 learned rules added, the deepest analyses of any model — and it still finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. Ironically, that’s a classic QA failure mode: comprehensive test coverage, poor release judgment. The same weakness, in weaker form, showed up in all four other models.
A Fairness Footnote Worth Reading
One caveat for the benchmark-literate: K3 ran without an effort parameter (the API default), while the other four models ran at xhigh. In other words, the newcomer may not even have been trying its hardest — which makes the 93 either more impressive or the comparison less apples-to-apples, depending on how you read it. Either way, the lesson stands.
You Can Play the Quiz Yourself
The results power a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html — worth a try if you think you can tell frontier models apart by their judgment rather than their prose.

Why This Matters for Builders
The live company behind the benchmark is no slide deck: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. It runs every business day, and you can watch it at firmulate.com, with full results and plain-language findings on the benchmarks page.
The takeaway for anyone shipping AI agents into a real stack: chat quality and management quality are different axes, and the league is wide open — a newcomer beat three of four Western frontier models. If an agent will touch your CRM, your support queue, or your forecast, “does it write well” is the wrong acceptance criterion. The right one is closer to what Firmulate measures: does it finish what it starts, does it read the files first, does it stay honest under pressure? Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
Picking a model without a test of your own isn’t a decision. It’s a bet. This experiment is the reminder — and the harness — to stop gambling.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI security and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
