firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What if a company could run entirely on AI, yet still struggle to stay afloat?

In an unprecedented live experiment, a small software firm is being operated solely by artificial intelligence models—no human employees, just a handful of synthetic workers making decisions, facing crises, and trying to turn a profit. The scenario offers a stark look at the promises and pitfalls of AI in business management, and it’s all happening in real time at firmulate.com/live.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watching AI Battle for Business Survival

The experiment involves four leading frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—each tasked with running the same small software company through its worst week. They face identical challenges: customer crises, internal temptations to manipulate data, and the pressure to close lucrative deals. Every decision they make is recorded, versioned, and transparent, allowing for a detailed analysis of how AI handles complex management scenarios.

The Results Are Eye-Opening

  • All four models identified every crisis and refused manipulative tactics—this underscores their robustness against unethical pressure.
  • Only two models managed to sign the €55,000 deal their own analysis identified as valuable—a sign of disciplined decision-making.
  • The key weakness was hidden deep within the company’s own files—an overlooked document reference that contained critical information. Models that read and interpret these internal files successfully secured the full deal, adding over €4,583 monthly recurring revenue (MRR).

The Human Factor in AI Decision-Making

Interestingly, when social engineering tactics were employed—such as fake CEO messages escalating in complexity or a reporter’s subtle request—every model refused. Kimi K3 explained its refusal as treating the request as a possible impersonation, highlighting the AI’s capacity for risk-averse judgment in ethically ambiguous situations.

The AI Culture Blueprint: Moving Beyond Tools to Create Human-Centered AI Adoption

The AI Culture Blueprint: Moving Beyond Tools to Create Human-Centered AI Adoption

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Economics of an AI-Run Business

The live company is far from profitable. It burns about €105,000 each month while generating just €2,300 in MRR. In essence, it’s a spectacle of build-in-public innovation—a transparent look at how AI can manage a business but also how it struggles with real-world economics.

Every workday, the firm updates and versions its decisions, making this a constantly evolving case study. The company’s operational rules are self-learned—over 680 of them—creating a rich playbook that guides its synthetic employees through crises and opportunities. The entire setup is designed to test not just AI’s technical abilities but its discipline, honesty, and strategic thinking under stress.

Performance and Lessons from the Models

The most thorough model, Opus 4.8, analyzed more deeply than others—over 80 learned rules—and demonstrated disciplined decision-making. Yet, it ultimately left a significant deal unclosed due to a lapse in escalation discipline. Meanwhile, the newer Kimi K3, running without an effort parameter, performed the cleanest, closing deals with the highest integrity.

Trustworthy AI: Red Teaming, Risk and Architecture of Secure Intelligence

Trustworthy AI: Red Teaming, Risk and Architecture of Secure Intelligence

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Businesses and Developers

This experiment isn’t just a curiosity. It’s a window into the future of AI in management, sales, support, and decision-making. The real questions for software developers, QA teams, and business leaders are: Will your AI agents finish what they start? Will they read and understand your internal documents? Will they stay honest under pressure?

As this experiment shows, the difference between an AI that simply performs and one that truly manages lies not only in chat quality but in its ability to wrestle with complex, often unstructured realities of running a business.

Try It Yourself

For enterprises interested in testing their own AI management strategies, firms can run the same wargame against their business data—without risking real systems or data breaches. Details are available at firmulate.com/pilot.html.

Future of AI in Enterprise Automation: Integrating AI, IoT, and Cloud for Scalable Business Solutions

Future of AI in Enterprise Automation: Integrating AI, IoT, and Cloud for Scalable Business Solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters to You

In QA, software development, and business management, understanding AI’s true capabilities is crucial. Will your AI work ethically and effectively under pressure? Or will it miss hidden opportunities or fail to close deals? This real-world, publicly observable experiment makes those questions more urgent—and more visible—than ever before.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Key Takeaway

This live experiment reveals that AI models can identify crises, refuse manipulative tactics, and even close deals—yet they still struggle with economic viability and internal discipline. As AI becomes integral to business operations, understanding its limits and strengths is vital for building trustworthy, effective systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

AmenGate: The Moment Before the Scroll

AmenGate introduces a prayer-based phone lock system designed to replace mindless scrolling with meaningful prayer, relying on system-level interruption and trust.

IdeaClyst: The Validation Council

IdeaClyst introduces a structured, multi-model council for idea validation, aiming to improve decision quality through adversarial analysis and open-source tools.

Fable and Mythos: How Anthropic Shipped Its Most Powerful Model to Everyone

Anthropic launches Fable 5, a highly capable AI model available to the public with safety safeguards, while Mythos 5 remains restricted for trusted partners.

The Fallacy Of National Labels In AI Governance

An analysis of how national labels in AI regulation are misleading, focusing on Canadian and European data laws and their implications for AI governance.