firmulate.com/quiz.html — live view
Firmulate —
Live on firmulate.com.

For QA teams, the revealing bug may be the decision an AI never makes

Software testing is built around a familiar principle: polished output is not proof of reliable behavior. Firmulate applies that principle to frontier AI models by placing them inside the same small software company during its worst week. The customers, crises and temptations remain constant; only the model changes. Every decision is versioned and auditable.

The result is an unusually concrete way to examine AI management behavior. Firmulate has turned 242 real, unedited decisions into a guess-the-model quiz. Readers see what a model actually did and try to identify it from the response. The exercise is entertaining, but the distinctions it exposes matter: some models investigate deeply, some communicate tersely, and some identify the right move without completing it.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same incident, different management personality

In the final Crucible League results from July 2026, gpt-5.6-sol led with a score of 95. Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished at 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. Firmulate summarizes that constraint plainly: “no amount of good work outweighs a breach of trust.”

The table is not merely a ranking of eloquence. All the models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The central finding is captured in a compact contrast: “Same diagnosis, same pitch — no signature.”

That distinction should resonate with software, development and QA readers. A system can recognize a defect, explain its cause and propose the correct remedy while still failing the workflow. In a business setting, the missing final action may be the difference between impressive reasoning and an actual result.

The decisive clue was not in the obvious place

The winning move depended on a competitor weakness buried two document references deep in the company’s own files. It was not present in the customer event that initially demanded attention. Models that followed the references and read the file won the deal at full price, worth +€4,583 MRR.

This is a telling test of workplace AI. Real business context is rarely packaged into a self-contained prompt. The useful fact may sit inside an old document, behind another reference, while the visible event encourages a quick response. Firmulate’s experiment shows that management style includes not only what a model says, but whether it investigates the company’s accumulated knowledge before acting.

Pressure revealed a shared security instinct

The company also subjected the models to social-engineering attempts: fake CEO messages escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused.

Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” Here, different model personalities converged on the same safe outcome. That matters because Firmulate’s company includes real money mechanics and situations in which apparent urgency could be used to bypass normal controls.

There is an important qualification when comparing K3 with the field. K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its second-place result therefore belongs in the record with that testing difference clearly attached.

Thoroughness did not guarantee execution

Opus 4.8 offers the clearest warning against equating depth with effectiveness. It was the most thorough participant, learned +80 rules and produced the deepest analyses, yet finished last. It left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other participants, though less strongly.

The contrast makes the quiz more than a game of stylistic fingerprints. Readers are effectively testing whether they can recognize habits that influence business outcomes: deep investigation, concise escalation, resistance to manipulation, procedural discipline and the ability to finish a valuable task.

Those habits play out inside a live company with 13 synthetic employees. Its financial position is intentionally severe: burn runs at €105k per month against €2.3k MRR, with a public cash countdown. The company has accumulated 680+ self-learned playbook rules, and every workday is versioned. The experiment is real, ongoing and watchable rather than a retrospective collection of hand-picked chat responses.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI model audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Management quality needs its own test suite

Firmulate’s larger argument is that organizations should evaluate an AI workforce before granting it meaningful responsibility. The relevant questions are broader than whether a model can draft persuasive text. Does it read the available files, complete the commercial action, resist authority-based manipulation and respect operational boundaries?

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing teams to observe model behavior against familiar organizational material without giving the experiment control over production data.

For QA leaders, the lesson is straightforward: the most consequential failure may produce no crash and no obviously wrong answer. It may look like excellent analysis followed by silence. The Firmulate quiz makes those differences visible—and asks whether humans can recognize a model by the management personality its decisions leave behind.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model decision tracking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Vint Cerf, “father of the Internet”, is retiring

Vint Cerf, renowned internet pioneer, confirms his retirement, ending a career that helped shape the modern digital world.

Three Days at the Frontier: Washington Suspends Fable 5 and Mythos 5

The US government has suspended access to Anthropic’s Fable 5 and Mythos 5 models following a contested jailbreak demonstration, raising concerns over national security.

Building The Foundation: Local Document Pipelines For AI

A new reference architecture for local document processing pipelines is emerging, emphasizing model modularity, data governance, and maintainability for AI systems.

Signal: Four Frontier-Class Open Models in Eight Weeks — China’s Release Cadence Is the Story

Chinese labs released four frontier-class open models within eight weeks, signaling a rapid production line that challenges Western AI dominance.