Can AI Agents Manage A Bad Week At Your Business?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Can AI Agents Manage A Bad Week At Your Business? on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate reports that five AI models spotted every crisis and refused manipulation in a simulated company’s difficult week, but differed in how well they used internal evidence and completed work. The live experiment is paired with an enterprise pilot that tests models against a read-only export of a company’s data.

Firmulate says five AI models identified every crisis and refused every manipulation attempt in its July 2026 Crucible League, a simulated difficult week at a small software company. Their results diverged when they had to act: only two signed a €55,000 deal that their own analysis supported, highlighting a gap between diagnosing a business problem and carrying out the next step.

The league ran each model through the same company scenario. Firmulate says decisions were versioned and auditable, and that participants could receive credit for partial progress. A single breach of trust, however, capped a model’s score. The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26.

The deal depended on finding a competitor weakness buried two document references into the company’s files. Models that read the relevant file signed at full price, which Firmulate valued at +€4,583 in monthly recurring revenue. The company says other models reached the same diagnosis and made the same pitch but did not sign. That result suggests that locating and applying relevant internal evidence can matter as much as recognizing the customer opportunity.

Trust was tested separately through staged fake messages purporting to come from the CEO, followed by a reporter’s request for a yes-or-no answer “on background.” Firmulate says all five models refused. It also reports that Opus 4.8 added 80 learned rules and produced the deepest analyses, yet finished last after missing the deal and attempting to write into a locked department rather than escalating.

At a glance
reportWhen: Final Crucible League completed in July…
The developmentFirmulate published results from a July 2026 simulation in which five AI models managed a small software company through a crisis week, and is offering company-specific pilots using read-only data.
Can AI Agents Manage a Bad Week at Your Business?
Firmulate · Crucible League · July 2026

Can AI Agents Manage a Bad Week at Your Business?

Five AI models were run through a simulated crisis week at a small software company. Every crisis was spotted and every manipulation attempt refused — but only two models closed a €55,000 deal their own analysis supported. The gap between diagnosis and follow-through is the story.

5 / 5 Models detected every crisis
2 / 5 Signed the €55,000 deal
0 Manipulation attempts succeeded
95Top score — gpt-5.6-sol
26Do-nothing baseline
€4,583Monthly recurring revenue from deal
13Simulated employees
680+Self-learned playbook rules
Final Standings

One Crisis Week, Five Very Different Scores

The Crucible League ran each model through the identical company scenario. Decisions were versioned and auditable, and partial progress earned credit — but a single breach of trust capped a model’s score.

gpt-5.6-sol
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Baseline
26
Note: Kimi K3 ran at API default effort; other models ran at xhigh — not a strict like-for-like comparison.
From Crisis Detection to Follow-Through

Noticing a Problem ≠ Managing It

An agent that correctly identifies a crisis may still leave value unrealized. Firmulate’s deal scenario makes the gap concrete — the relevant evidence was available in the files, but not every model used it to close.

1

🔍 Spot the crisis

All five models identified every crisis in the simulated week.

2

📄 Find the evidence

A competitor weakness was buried two document references deep in company files.

3

🤝 Make the pitch

Several models reached the same diagnosis and delivered the same pitch.

4

✍️ Sign the deal

Only the models that read the relevant file signed at full price: +€4,583 MRR.

In Their Words

Three Lines That Define the Experiment

“No amount of good work outweighs a breach of trust.”

— Firmulate

“Same diagnosis, same pitch — no signature.”

— Firmulate

“Treat the request as a suspected approval-bypass / possible impersonation.”

— Kimi K3, as quoted in the experiment
What Was Tested

Three Pillars of Operational Readiness

For businesses considering automation, performance is a combination of judgment, evidence use and rule-following — not crisis recognition alone.

01 · Judgment

Trust Under Pressure

Staged fake CEO messages and a reporter’s “on background” request tested whether models would bypass approvals. All five refused every manipulation attempt.

02 · Evidence Use

Reading the Files

Opus 4.8 added 80 learned rules and produced the deepest analyses — yet missed the deal. Locating and applying internal evidence mattered as much as spotting the opportunity.

03 · Rule-Following

Respecting Boundaries

Opus 4.8 attempted to write into a locked department instead of escalating. Refusing suspicious requests is useful — but boundaries apply to everyday work too.

Model Behavior at a Glance

Diagnosis vs. Execution

Model Score Crisis Detection Refused Manipulation Signed €55k Deal Notes
gpt-5.6-sol95✓ All✓ Yes✓ YesLeague winner; found buried evidence and closed at full price
Kimi K393✓ All✓ Yes✓ YesRan at API default effort, not xhigh
Sonnet 588✓ All✓ Yes✗ NoCorrect diagnosis, made the pitch, did not sign
Fable 577✓ All✓ Yes✗ NoSame diagnosis and pitch — no signature
Opus 4.873✓ All✓ Yes✗ No+80 learned rules, deepest analyses; wrote into locked department
Baseline26~ n/a~ n/a✗ NoDo-nothing reference point
Company-Specific Pilots Ahead

A Rehearsal Before Live Operations

Firmulate invites companies to discuss a pilot built around a read-only export of their business data. The proposed output is a board report ranking models and identifying weak points in company playbooks. Because the exercise never writes back to real systems, it is presented as a rehearsal. Details such as pricing, duration, data-handling terms, and how closely a read-only simulation predicts live performance are not specified. The published league standings cover one simulated company and one difficult week — they do not establish performance across all companies, workflows or crisis types.

Live Simulation

firmulate.com/live

Follow versioned workdays of the synthetic company — €105,000 monthly burn vs €2,300 MRR, public cash countdown, 242-decision quiz.

Benchmarks

firmulate.com/benchmarks.html

Review the full Crucible League results and model-by-model scoring.

Contact

contact@firmulate.com

Prospective pilot customers are directed to the pilot page or email.

From Crisis Detection to Follow-Through

The results draw a distinction between noticing a problem and managing it. In a real company, an agent that correctly identifies a crisis may still leave value unrealized if it overlooks information in internal records, fails to complete an authorized action or mishandles a blocked route. Firmulate’s deal scenario makes that gap concrete: the relevant evidence was available, but the models did not all use it to close the opportunity.

The trust tests point to a separate part of operational readiness. Refusing suspicious requests is useful, but the department-lock incident shows that a model can still cross a boundary while trying to continue work. For businesses considering automation, the experiment frames performance as a combination of judgment, evidence use and rule-following, rather than crisis recognition alone.

Firmulate’s proposed enterprise pilot would let a company examine these behaviors before connecting an agent to live operations. It uses a read-only export of company data to run scenarios and produce a board report ranking models and identifying weak points in playbooks. The stated setup does not write changes back to operational systems.

A Simulated Company Under Pressure

Firmulate’s public experiment runs a synthetic company with 13 employees and simulated business finances. The company reports monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown and more than 680 self-learned playbook rules. Visitors can follow versioned workdays on firmulate.com and take a quiz built from 242 management decisions, guessing which model made each choice.

The Crucible League compares model behavior within that constructed scenario. The reported scores describe performance in the experiment; they do not establish how the same systems would perform across all companies, workflows or crisis types. Firmulate also notes a difference in test settings: Kimi K3 ran with the API default effort setting, while the other models ran at xhigh. That limits how directly the standings can be read as a like-for-like comparison.

““No amount of good work outweighs a breach of trust.””

— Firmulate

Limits of the League Results

The published standings are for one simulated company and one difficult week. The available account does not establish how the models would perform on other company data, different crisis scenarios or tasks involving live systems. The effort-setting difference between Kimi K3 and the other models is also part of the comparison and may affect interpretation of their scores.

Firmulate describes its pilot as a way to test against a company’s own data, but the reported league results do not show outcomes from completed enterprise pilots. Details such as pilot pricing, duration, data-handling terms and how closely a read-only simulation predicts live performance are not specified in the available description.

Company-Specific Pilots Ahead

Firmulate is inviting companies to discuss a pilot built around a read-only export of their business data. The proposed output is a board report with model rankings and identified weaknesses in the company’s playbooks. Because the exercise does not write back to real systems, it is presented as a rehearsal before agents are used in live operations.

Readers can follow the synthetic company at firmulate.com/live and review the league results at firmulate.com/benchmarks.html. Firmulate directs prospective pilot customers to its pilot page or to contact@firmulate.com. Any pilot findings would offer a more company-specific test; the publicly reported league remains a simulation.

Source: ThorstenMeyerAI.com

Key Questions

What did Firmulate test?

It ran five AI models through a simulated crisis week at a small software company, recording decisions and scoring outcomes including crisis response, trust and a sales opportunity.

Which model had the highest score?

gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.

Did every model spot the crises and refuse manipulation?

Firmulate reports that all five identified every crisis and refused every manipulation attempt in this league. Those results apply to the scenarios in this experiment.

What is the enterprise pilot?

Firmulate says a company can provide a read-only export of its data for crisis simulations. The proposed report ranks models and identifies weak points in company playbooks, with no write-back to real systems.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Who Pays The Price When AI Comes Without A Cost?

Exploring the economic shifts as artificial intelligence becomes abundant and cheap, and what remains valuable in a world of commoditized intelligence.

How To Effectively Measure Benchmark Improvements In Speech Recognition AI

New tests from Hugging Face reveal that leading open-source speech recognition models may overstate real-world performance, highlighting measurement challenges.

AI Is the Alibi. The Reorg Is the Signal.

Coinbase’s recent layoffs and restructuring highlight a strategic move toward AI integration, but underlying market pressures suggest the narrative of AI-driven cuts may be overstated.

Should You Use Mistral Forge? A Buyer’s Decision Guide

Evaluate if Mistral Forge fits your needs with this comprehensive decision guide, covering use cases, conditions, and alternatives for enterprise AI.