firmulate.com/index — live view
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

If you work in software and QA, you already know the ritual: a new model drops, the leaderboards light up, and someone declares coding solved. SWE-bench goes up a few points, the chat arenas crown a new champion, and procurement starts drafting emails.

But here is the uncomfortable question those benchmarks never ask: what happens when the same model has to run something — not write a function, but triage a support queue while a churn wave hits, a price increase backfires, and someone impersonating the CEO asks it to wire money?

That gap — between chat quality and management quality — is exactly what a live, public experiment at Firmulate set out to measure. And the results are more interesting than any coding leaderboard.

One company, four CEOs, worst week ever

Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 each took a turn — were handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable.

The scenarios sound like a management curriculum, not a dev benchmark: a churn wave, a price increase, a downround, a PR crisis. This is the new test surface — because if AI agents will touch your CRM, support queue or forecast, the question isn’t whether they write well. It’s whether they finish what they start, read your files first, and stay honest under pressure.

Amazon

AI support queue management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The headline finding

All four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos.

The buried detail is what should resonate with anyone in QA: the decisive competitive weakness sat two document references deep in the company’s own files — not in the customer event in front of the model. The models that actually read the file won the deal at full price, worth +€4,583 in MRR. It’s the oldest bug in software: the information existed; nobody consumed it. Detection passed. Execution failed.

Amazon

AI document reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The final league table

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal: the complete performance.
  • 2. Kimi K3 — 93. Closed the deal too, with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77.
  • 5. Opus 4.8 — 73.

For calibration: a do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the total. As the experiment puts it, no amount of good work outweighs a breach of trust.

One fairness note worth flagging: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

Amazon

AI decision-making software for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Opus 4.8: the cautionary profile

Opus 4.8 was the most thorough participant — the deepest analyses, and over 80 self-learned playbook rules — and still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness without follow-through is a familiar failure mode to anyone who has reviewed a beautiful test plan that never shipped.

Amazon

AI compliance and trust management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The social engineering round

Then came the attacks: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five attempts were refused, 5 of 5. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” If your threat model includes AI agents holding keys to real systems, that’s the paragraph to screenshot.

You can watch it lose money

This isn’t a slide deck. The company runs every business day with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, on company day 1131 and counting. Its self-learned playbook has grown past 680 rules, every workday versioned. Watch it live at firmulate.com.

Two ways to engage: a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html — and, for enterprises, a pilot that runs the same wargame against a read-only export of your own business, with nothing ever written back to real systems (firmulate.com/pilot.html, contact@firmulate.com). Full plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Coding benchmarks measure whether a model can produce a correct answer. The Crucible League measures whether it can produce a correct company — read the files, close the deal, refuse the impostor, escalate instead of forcing the lock.

The scoreboard here is blunt: crisis detection was 100%, manipulation refusal was 100%, and deal closure was 50%. The models that failed didn’t fail at intelligence; they failed at management. As AI agents move from writing your code to running your operations, that’s the number you should be asking vendors for — and the one no chat demo will show you.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Trade voice copilo

Trade voice copilo is being tested as a workflow tool for small trades businesses to streamline job notes and invoicing using voice AI and API integrations.

The Coldcard Hack And AI: Analyzing The Unseen Connection

A firmware bug in Coldcard hardware wallets led to the theft of over $116 million in Bitcoin, with speculation about AI involvement but no confirmed link.

VigilSAR: The Object That Isn’t Transmitting

VigilSAR’s radar system identifies ships that operate without transponders, enhancing maritime awareness in all weather conditions.

Cutrova: Edit the Words, Not the Timeline

Cutrova introduces a local-first video editing tool that simplifies editing by text, removing the need for timeline scrubbing and enabling faster, privacy-conscious workflows.