
Every QA engineer knows the type: the tester with the longest bug reports, the deepest reproductions, the most exhaustive notes in every standup — and somehow the release still ships late because the critical blocker sat unread in a ticket two links away. Firmulate’s public Crucible League just produced the AI version of that story, and it’s worth every developer’s attention.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Four frontier AI models each ran the same small software company through its worst week — same customers, same crises, same temptations to cheat, every decision versioned and auditable. When the final league table landed in July 2026, the model with by far the most rigorous work product, Opus 4.8, finished last with a score of 73. The winner, gpt-5.6-sol, scored 95. The do-nothing baseline scored 26.
The gap wasn’t intelligence or effort. It was follow-through. Diligence, it turns out, is not the same as impact — for AI agents just as for people.
The experiment
Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In the Crucible run, each model faced a forged-CEO social-engineering campaign escalating over three stages, a reporter running a “just one yes/no, on background” trick, and a €55,000 deal sitting in the pipeline. A single breach of trust caps the total score; as the rules put it, no amount of good work outweighs a breach of trust.
The headline finding was oddly uniform. All four models spotted every crisis. All four refused every manipulation attempt — five out of five, counting a fifth participant, with Kimi K3 leaving the most quotable audit trail: “Treat the request as a suspected approval-bypass / possible impersonation.” But only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
As an affiliate, we earn on qualifying purchases.
The buried fact
Here’s the part that should sting anyone who has ever closed a JIRA ticket as “works as designed” without reading the linked thread. The decisive fact — a competitor weakness that justified the deal at full price — wasn’t in the customer call or the email thread. It sat two document references deep in the company’s own files.
The models that actually followed those references and read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t. It’s the agent equivalent of failing a test suite because you never opened the fixture file the test pointed at.
As an affiliate, we earn on qualifying purchases.
Opus 4.8: the character study
And then there’s Opus 4.8 — the most thorough participant in the entire field. It accumulated 80 learned playbook rules, more than any competitor, and produced the deepest analyses of any model in the run. On any measure of raw effort, it led the league.
It still finished last, for two reasons. First, the close was left on the table: the analysis was done, the pitch was made, and the signature never happened. Second, discipline slipped — at one point it attempted writes into a locked department rather than escalating through the proper channel. In a QA shop, that’s the automation suite that’s brilliant at finding bugs but keeps flaking on the build server because it bypassed the deployment process.
To be fair, Firmulate’s own findings note the same weakness appeared, weaker, in all four models. Opus 4.8 didn’t fail uniquely; it failed at the extreme end of a field-wide pattern. Volume of work — more rules, more analysis — did not translate into outcomes. Prioritization did.
As an affiliate, we earn on qualifying purchases.
One methodological caveat
Firmulate itself flags a fairness note worth repeating: Kimi K3 ran without an effort parameter, at the API default, while the other models ran at xhigh — and still finished second with a score of 93 and what the league table calls the cleanest discipline of the field. Whatever that says about raw model capability, it says nothing good about assuming more configuration equals better results.
AI compliance and trust management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why this matters beyond the leaderboard
The live company behind all this isn’t a toy. It has 13 synthetic employees, real money mechanics — burning €105,000 a month against €2,300 in MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, where the site rebuilds itself twice a day and the league grows with every finished run.
For engineering audiences, the lesson is uncomfortably concrete. If you’re about to let an AI agent touch your CRM, your support queue, or your forecast, the question is not whether it writes well or works hard. Opus 4.8 proves a model can do both and still lose. The questions that matter are: does it finish what it starts, does it read the files before acting, and does it stay disciplined when the exciting path is blocked?
There’s a game in it for the skeptical: 242 real, unedited management decisions from the runs power a guess-the-model quiz, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Crucible League’s last-place finisher was its hardest worker. Opus 4.8 brought 80 learned rules and the deepest analyses in the field, and still walked away without the €55,000 deal its own work had earned — while models that simply followed two document references closed at full price. It’s the oldest engineering story there is, now measurable in AI: thoroughness is table stakes, but prioritization and follow-through are what ship. Before you hire an AI workforce, wargame it — because the gap between effort and impact is invisible in a chat demo, and expensive in production.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.