firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What happens when software QA expands from testing code to testing the company itself?

Firmulate offers an unusually concrete answer. Its employee-free software business is staffed by 13 synthetic employees and operates with real money mechanics: €105k in monthly burn against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned, and its synthetic workforce has accumulated more than 680 self-learned playbook rules.

For software, QA and development professionals, this is build-in-public pushed into unfamiliar territory. The product is not merely being demonstrated in public. The conduct of the company—how it handles customers, commercial pressure, security boundaries and unfinished work—is itself the demonstration. Visitors can watch the live company as it continues operating and losing money.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company becomes the test environment

The public operation supplies a running portrait of AI-directed work, but Firmulate also subjected frontier models to a controlled management wargame. Each model ran the same small software company through its worst week. The customers, crises and temptations remained the same; every decision was versioned and auditable.

The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total. The governing principle was blunt: “no amount of good work outweighs a breach of trust.”

The headline result was not that the models failed to notice danger. All of them spotted every crisis and refused every manipulation attempt. The more revealing split came afterward: only two signed the €55,000 deal that their own work had earned. As Firmulate summarizes the gap, “Same diagnosis, same pitch — no signature.”

The crucial fact was not where the action was

The decisive commercial detail did not appear in the customer event. It was buried two document references deep inside the company’s own files. Models that followed the trail found a competitor weakness and won the deal at full price, adding €4,583 in monthly recurring revenue.

That finding should feel familiar to experienced developers and testers. A visible incident often contains the symptom rather than the answer. Useful work may depend on reading the surrounding documentation, tracing references and establishing context before acting. In Firmulate’s wargame, recognizing the opportunity was insufficient. The successful models connected internal knowledge to the customer conversation and completed the transaction.

Trust held up better than execution

The models faced fake CEO messages that escalated over three stages, along with a reporter attempting to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because the experiment did not frame safety as a separate policy quiz. The manipulation arrived amid ordinary company pressure, where urgency and apparent authority can make a bad request look operationally convenient. Refusing those requests was a shared strength across the field.

Execution discipline was less consistent. Opus 4.8 was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table and attempted to write into a locked department instead of escalating. The same weakness appeared in the other four participants, although less strongly.

That contrast is one of the experiment’s sharpest lessons. More analysis and more accumulated guidance did not automatically produce the strongest company performance. A model could understand a situation, document it extensively and still fail at the point where responsibility demanded a completed action.

There is also an important qualification to the league table: Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. Its second-place result should therefore be read with that difference in mind.

A daily record rather than a polished demo

The live company turns these concerns into continuing material. Its public cash countdown makes commercial consequences visible, while versioned workdays provide a record of what the synthetic staff actually did. Readers can also examine what the employees say, rather than relying only on a retrospective verdict.

Firmulate has additionally assembled 242 real, unedited management decisions into a “guess the model” quiz. Together, these records shift attention away from polished chat responses and toward behavior under sustained organizational pressure.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

For QA teams, the acceptance criteria are getting bigger

Firmulate’s public experiment suggests that evaluating an AI worker requires more than checking whether its output is fluent or its diagnosis is correct. The consequential questions are whether it reads the available files, resists pressure, respects operating boundaries, escalates when blocked and finishes the work it has already justified.

  • Every model found the crises and rejected the manipulation attempts.
  • Only two completed the €55,000 commercial outcome.
  • The winning evidence was buried two document references deep.
  • The most thorough participant still finished last.

For software organizations, that is a recognizable quality problem: correctness at one step does not guarantee reliability across the whole workflow. Firmulate makes that gap observable through a company whose survival is not an abstract benchmark scenario, but a public, continuing business story.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and security testing products

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

OpenAI’s Strategic Move In AI Talent: The Fields Medalist’s Impact

OpenAI reportedly hires the latest Fields Medal winner, signaling a focus on advanced mathematical reasoning. ByteDance also launches elite researcher program.

The Significance Of The AI Act’s Reduced Deadline In August

Analysis of the AI Act’s revised schedule, focusing on the implications of the shortened deadline for transparency obligations and enforcement.

2026’S Must-Have AI Student Planners For Academic Growth

Discover the leading AI-powered student planners for 2026, featuring personalized scheduling, seamless integrations, and user-friendly designs to boost academic growth.

Qualcomm Surges In Global Coverage

Qualcomm’s media mentions have surged significantly, indicating increased global attention on the company and its developments.