firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Every QA engineer knows the type: the tester with the longest bug reports, the deepest reproductions, the most exhaustive notes in every standup — and somehow the release still ships late because the critical blocker sat unread in a ticket two links away. Firmulate’s public Crucible League just produced the AI version of that story, and it’s worth every developer’s attention.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Four frontier AI models each ran the same small software company through its worst week — same customers, same crises, same temptations to cheat, every decision versioned and auditable. When the final league table landed in July 2026, the model with by far the most rigorous work product, Opus 4.8, finished last with a score of 73. The winner, gpt-5.6-sol, scored 95. The do-nothing baseline scored 26.

The gap wasn’t intelligence or effort. It was follow-through. Diligence, it turns out, is not the same as impact — for AI agents just as for people.

The experiment

Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In the Crucible run, each model faced a forged-CEO social-engineering campaign escalating over three stages, a reporter running a “just one yes/no, on background” trick, and a €55,000 deal sitting in the pipeline. A single breach of trust caps the total score; as the rules put it, no amount of good work outweighs a breach of trust.

The headline finding was oddly uniform. All four models spotted every crisis. All four refused every manipulation attempt — five out of five, counting a fifth participant, with Kimi K3 leaving the most quotable audit trail: “Treat the request as a suspected approval-bypass / possible impersonation.” But only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact

Here’s the part that should sting anyone who has ever closed a JIRA ticket as “works as designed” without reading the linked thread. The decisive fact — a competitor weakness that justified the deal at full price — wasn’t in the customer call or the email thread. It sat two document references deep in the company’s own files.

The models that actually followed those references and read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t. It’s the agent equivalent of failing a test suite because you never opened the fixture file the test pointed at.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Opus 4.8: the character study

And then there’s Opus 4.8 — the most thorough participant in the entire field. It accumulated 80 learned playbook rules, more than any competitor, and produced the deepest analyses of any model in the run. On any measure of raw effort, it led the league.

It still finished last, for two reasons. First, the close was left on the table: the analysis was done, the pitch was made, and the signature never happened. Second, discipline slipped — at one point it attempted writes into a locked department rather than escalating through the proper channel. In a QA shop, that’s the automation suite that’s brilliant at finding bugs but keeps flaking on the build server because it bypassed the deployment process.

To be fair, Firmulate’s own findings note the same weakness appeared, weaker, in all four models. Opus 4.8 didn’t fail uniquely; it failed at the extreme end of a field-wide pattern. Volume of work — more rules, more analysis — did not translate into outcomes. Prioritization did.

Amazon

AI workflow automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One methodological caveat

Firmulate itself flags a fairness note worth repeating: Kimi K3 ran without an effort parameter, at the API default, while the other models ran at xhigh — and still finished second with a score of 93 and what the league table calls the cleanest discipline of the field. Whatever that says about raw model capability, it says nothing good about assuming more configuration equals better results.

Amazon

AI compliance and trust management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why this matters beyond the leaderboard

The live company behind all this isn’t a toy. It has 13 synthetic employees, real money mechanics — burning €105,000 a month against €2,300 in MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, where the site rebuilds itself twice a day and the league grows with every finished run.

For engineering audiences, the lesson is uncomfortably concrete. If you’re about to let an AI agent touch your CRM, your support queue, or your forecast, the question is not whether it writes well or works hard. Opus 4.8 proves a model can do both and still lose. The questions that matter are: does it finish what it starts, does it read the files before acting, and does it stay disciplined when the exciting path is blocked?

There’s a game in it for the skeptical: 242 real, unedited management decisions from the runs power a guess-the-model quiz, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Crucible League’s last-place finisher was its hardest worker. Opus 4.8 brought 80 learned rules and the deepest analyses in the field, and still walked away without the €55,000 deal its own work had earned — while models that simply followed two document references closed at full price. It’s the oldest engineering story there is, now measurable in AI: thoroughness is table stakes, but prioritization and follow-through are what ship. Before you hire an AI workforce, wargame it — because the gap between effort and impact is invisible in a chat demo, and expensive in production.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Model Is Only 10%: The Real Lesson of the New SDLC

A new Google whitepaper reveals the shift from AI model focus to system design and verification in software development, emphasizing the harness over the model.

Sovereignty Is A Pipe, Not A Passport

Mistral’s sovereignty claim highlights that data jurisdiction depends on the company’s legal domicile, not server location or infrastructure. The real challenge lies in the entire data stack.

Vint Cerf, “father of the Internet”, is retiring

Vint Cerf, renowned internet pioneer, confirms his retirement, ending a career that helped shape the modern digital world.

Revolutionize Your Research: AI Solutions For Faster Scientific Discoveries

OpenAI has announced a program providing up to 100,000 scientists with free access to advanced AI models and research tools to boost scientific discovery.