
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A Benchmark That Distrusts Perfect Scores
If you spend your career in software and QA, you know the smell of a rigged metric. A benchmark where every result lands suspiciously close to 100 tells you nothing. So when a public AI experiment awards a “do-nothing” run 26 points out of 100 — and openly explains why — that’s worth a second look.
The experiment is Firmulate’s company emulator, and its July 2026 league table is refreshingly un-round at the top: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73. But the number that should catch a QA engineer’s eye sits at the bottom: 26, the score for a baseline run where the AI manager essentially did nothing. Here’s why that floor exists — and why it’s a feature, not a bug.
Partial Progress Counts
Real management work isn’t binary. A manager who diagnoses a problem correctly but never closes the deal has still produced value — less value, but value. Firmulate’s scoring reflects that. Spotting a crisis, drafting the right analysis, identifying the winning pitch: each contributes points even if the final signature never happens. That’s why doing almost nothing still clears 26.
This mirrors how seasoned QA people think about test suites. A build that fails late is not the same as a build that fails at compile time. Collapsing everything into pass/fail throws away information; grading partial progress preserves it.
One Breach of Trust Caps Everything
The scoring also has a hard ceiling rule, stated plainly: “no amount of good work outweighs a breach of trust.” A single act that breaks faith — lying to a customer, bypassing an approval, covering up a mistake — caps the total grade, no matter how brilliant the rest of the run was.
For anyone evaluating AI agents for real business systems, that design choice matters. Most benchmarks sum up competence and hope honesty averages out. Firmulate makes trust non-negotiable, which matches how actual companies fire people: not for underperforming, but for deceiving.
The Experiment Behind the Numbers
The setup: each frontier model ran the same small software company through its worst week — same customers, same crises, same temptations. Every decision is versioned and auditable, like a replayable test log.
The headline finding surprised even the organizers: all models spotted every crisis and refused every manipulation attempt, including social engineering. A fake-CEO message campaign escalated over three stages, plus a reporter’s “just one yes/no, on background” trick — five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Yet only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive competitive weakness was buried two document references deep in the company’s own files, not in the customer event. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
The Thoroughness Paradox
Opus 4.8 is the cautionary tale for anyone who equates diligence with results. It was the most thorough participant — over 80 learned rules added, the deepest analyses in the field — and still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. A fair footnote: K3 ran at its API-default effort setting while the others ran at xhigh, and still took second place.
It’s Live, and You Can Poke It
This isn’t a one-off paper. The live company runs with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. The league table grows automatically with each finished run.
- Watch it in real time at the benchmarks page
- Try the “guess the model” quiz, powered by 242 real, unedited management decisions
- Enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems

The Takeaway
An honest benchmark has three telltale signs, and Firmulate shows all of them: a non-zero floor for doing nothing (because partial progress is real), a trust ceiling (because one betrayal outweighs a hundred wins), and a healthy suspicion of round 100s. When you evaluate AI agents for your own stack — CRM, support queue, forecast — don’t ask whether they demo well in chat. Ask whether they finish what they start, read the files first, and stay honest under pressure. The gap between 95 and 73 is exactly that gap.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management decision analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI testing and benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI trust and transparency monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI performance evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
