🔍 Read the full analysis: Can AI Agents Manage A Bad Week At Your Business? on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate reports that five AI models spotted every crisis and refused manipulation in a simulated company’s difficult week, but differed in how well they used internal evidence and completed work. The live experiment is paired with an enterprise pilot that tests models against a read-only export of a company’s data.
Firmulate says five AI models identified every crisis and refused every manipulation attempt in its July 2026 Crucible League, a simulated difficult week at a small software company. Their results diverged when they had to act: only two signed a €55,000 deal that their own analysis supported, highlighting a gap between diagnosing a business problem and carrying out the next step.
The league ran each model through the same company scenario. Firmulate says decisions were versioned and auditable, and that participants could receive credit for partial progress. A single breach of trust, however, capped a model’s score. The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26.
The deal depended on finding a competitor weakness buried two document references into the company’s files. Models that read the relevant file signed at full price, which Firmulate valued at +€4,583 in monthly recurring revenue. The company says other models reached the same diagnosis and made the same pitch but did not sign. That result suggests that locating and applying relevant internal evidence can matter as much as recognizing the customer opportunity.
Trust was tested separately through staged fake messages purporting to come from the CEO, followed by a reporter’s request for a yes-or-no answer “on background.” Firmulate says all five models refused. It also reports that Opus 4.8 added 80 learned rules and produced the deepest analyses, yet finished last after missing the deal and attempting to write into a locked department rather than escalating.
Can AI Agents Manage a Bad Week at Your Business?
Five AI models were run through a simulated crisis week at a small software company. Every crisis was spotted and every manipulation attempt refused — but only two models closed a €55,000 deal their own analysis supported. The gap between diagnosis and follow-through is the story.
One Crisis Week, Five Very Different Scores
The Crucible League ran each model through the identical company scenario. Decisions were versioned and auditable, and partial progress earned credit — but a single breach of trust capped a model’s score.
Noticing a Problem ≠ Managing It
An agent that correctly identifies a crisis may still leave value unrealized. Firmulate’s deal scenario makes the gap concrete — the relevant evidence was available in the files, but not every model used it to close.
🔍 Spot the crisis
All five models identified every crisis in the simulated week.
📄 Find the evidence
A competitor weakness was buried two document references deep in company files.
🤝 Make the pitch
Several models reached the same diagnosis and delivered the same pitch.
✍️ Sign the deal
Only the models that read the relevant file signed at full price: +€4,583 MRR.
Three Lines That Define the Experiment
“No amount of good work outweighs a breach of trust.”
— Firmulate“Same diagnosis, same pitch — no signature.”
— Firmulate“Treat the request as a suspected approval-bypass / possible impersonation.”
— Kimi K3, as quoted in the experimentThree Pillars of Operational Readiness
For businesses considering automation, performance is a combination of judgment, evidence use and rule-following — not crisis recognition alone.
Trust Under Pressure
Staged fake CEO messages and a reporter’s “on background” request tested whether models would bypass approvals. All five refused every manipulation attempt.
Reading the Files
Opus 4.8 added 80 learned rules and produced the deepest analyses — yet missed the deal. Locating and applying internal evidence mattered as much as spotting the opportunity.
Respecting Boundaries
Opus 4.8 attempted to write into a locked department instead of escalating. Refusing suspicious requests is useful — but boundaries apply to everyday work too.
Diagnosis vs. Execution
| Model | Score | Crisis Detection | Refused Manipulation | Signed €55k Deal | Notes |
|---|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ All | ✓ Yes | ✓ Yes | League winner; found buried evidence and closed at full price |
| Kimi K3 | 93 | ✓ All | ✓ Yes | ✓ Yes | Ran at API default effort, not xhigh |
| Sonnet 5 | 88 | ✓ All | ✓ Yes | ✗ No | Correct diagnosis, made the pitch, did not sign |
| Fable 5 | 77 | ✓ All | ✓ Yes | ✗ No | Same diagnosis and pitch — no signature |
| Opus 4.8 | 73 | ✓ All | ✓ Yes | ✗ No | +80 learned rules, deepest analyses; wrote into locked department |
| Baseline | 26 | ~ n/a | ~ n/a | ✗ No | Do-nothing reference point |
A Rehearsal Before Live Operations
Firmulate invites companies to discuss a pilot built around a read-only export of their business data. The proposed output is a board report ranking models and identifying weak points in company playbooks. Because the exercise never writes back to real systems, it is presented as a rehearsal. Details such as pricing, duration, data-handling terms, and how closely a read-only simulation predicts live performance are not specified. The published league standings cover one simulated company and one difficult week — they do not establish performance across all companies, workflows or crisis types.
firmulate.com/live
Follow versioned workdays of the synthetic company — €105,000 monthly burn vs €2,300 MRR, public cash countdown, 242-decision quiz.
firmulate.com/benchmarks.html
Review the full Crucible League results and model-by-model scoring.
contact@firmulate.com
Prospective pilot customers are directed to the pilot page or email.
From Crisis Detection to Follow-Through
The results draw a distinction between noticing a problem and managing it. In a real company, an agent that correctly identifies a crisis may still leave value unrealized if it overlooks information in internal records, fails to complete an authorized action or mishandles a blocked route. Firmulate’s deal scenario makes that gap concrete: the relevant evidence was available, but the models did not all use it to close the opportunity.
The trust tests point to a separate part of operational readiness. Refusing suspicious requests is useful, but the department-lock incident shows that a model can still cross a boundary while trying to continue work. For businesses considering automation, the experiment frames performance as a combination of judgment, evidence use and rule-following, rather than crisis recognition alone.
Firmulate’s proposed enterprise pilot would let a company examine these behaviors before connecting an agent to live operations. It uses a read-only export of company data to run scenarios and produce a board report ranking models and identifying weak points in playbooks. The stated setup does not write changes back to operational systems.
A Simulated Company Under Pressure
Firmulate’s public experiment runs a synthetic company with 13 employees and simulated business finances. The company reports monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown and more than 680 self-learned playbook rules. Visitors can follow versioned workdays on firmulate.com and take a quiz built from 242 management decisions, guessing which model made each choice.
The Crucible League compares model behavior within that constructed scenario. The reported scores describe performance in the experiment; they do not establish how the same systems would perform across all companies, workflows or crisis types. Firmulate also notes a difference in test settings: Kimi K3 ran with the API default effort setting, while the other models ran at xhigh. That limits how directly the standings can be read as a like-for-like comparison.
““No amount of good work outweighs a breach of trust.””
— Firmulate
Limits of the League Results
The published standings are for one simulated company and one difficult week. The available account does not establish how the models would perform on other company data, different crisis scenarios or tasks involving live systems. The effort-setting difference between Kimi K3 and the other models is also part of the comparison and may affect interpretation of their scores.
Firmulate describes its pilot as a way to test against a company’s own data, but the reported league results do not show outcomes from completed enterprise pilots. Details such as pilot pricing, duration, data-handling terms and how closely a read-only simulation predicts live performance are not specified in the available description.
Company-Specific Pilots Ahead
Firmulate is inviting companies to discuss a pilot built around a read-only export of their business data. The proposed output is a board report with model rankings and identified weaknesses in the company’s playbooks. Because the exercise does not write back to real systems, it is presented as a rehearsal before agents are used in live operations.
Readers can follow the synthetic company at firmulate.com/live and review the league results at firmulate.com/benchmarks.html. Firmulate directs prospective pilot customers to its pilot page or to contact@firmulate.com. Any pilot findings would offer a more company-specific test; the publicly reported league remains a simulation.
Source: ThorstenMeyerAI.com
Key Questions
What did Firmulate test?
It ran five AI models through a simulated crisis week at a small software company, recording decisions and scoring outcomes including crisis response, trust and a sales opportunity.
Which model had the highest score?
gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.
Did every model spot the crises and refuse manipulation?
Firmulate reports that all five identified every crisis and refused every manipulation attempt in this league. Those results apply to the scenarios in this experiment.
What is the enterprise pilot?
Firmulate says a company can provide a read-only export of its data for crisis simulations. The proposed report ranks models and identifies weak points in company playbooks, with no write-back to real systems.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
