🔍 Read the full analysis: How A Fixed Benchmark Ensures AI Managers Don’t Score Zero on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A novel fixed benchmark evaluates AI management performance during a simulated worst-week scenario. It assigns a minimum score of 26 for partial work and caps trust breaches, preventing zero scores and promoting accountability. This approach highlights the importance of trust and completion in AI-driven business management.
A new benchmark league developed by Firmulate has introduced a fixed scoring system for evaluating AI managers during a simulated week of business crises. The system ensures that no AI management score can be zero by assigning a minimum of 26 points for partial work, even when models do nothing, and caps trust breaches at a maximum score. This approach aims to promote accountability and realistic assessment of AI management capabilities.
The benchmark involved four frontier AI models managing a small software company through seven days of crises, customer manipulations, and trust attacks. Each model’s decisions were fully auditable, and scores reflected both their ability to handle crises and maintain trust. The top performer, gpt-5.6-sol, scored 95, while the baseline that did nothing scored 26, illustrating that partial progress is recognized but zero is avoided.
The scoring system’s core principle is that even minimal work is worth some points, but breaches of trust, such as failing to escalate or follow rules, limit the maximum achievable score. For more details, see the original analysis on Thorsten Meyer AI. The benchmark’s designers explicitly designed the scale to prevent artificially inflated perfect scores, with 100 being considered suspicious and potentially unmeasured. This structure encourages models to demonstrate integrity alongside competence, as detailed in the original analysis.
Key findings include that models which read their documentation and refused manipulation attempts performed better, winning deals worth €4,583 monthly recurring revenue. Conversely, thorough analysis alone did not guarantee success if models failed to follow through or escalate properly. The results underscore the importance of discipline, trust, and thoroughness in AI management tasks.
How A Fixed Benchmark Ensures AI Managers Don’t Score Zero
Firmulate’s benchmark league puts four frontier AI models in charge of a small software company for seven days of crises, customer manipulations, and trust attacks — with fixed scoring rules that guarantee no score of zero and no artificially perfect one.
Partial Work Counts. Trust Breaches Cap the Ceiling.
The scale is deliberately engineered: a floor of 26 points recognizes minimal but necessary management activity such as triaging inboxes and reading documentation. A single trust breach — failing to escalate, breaking rules — hard-limits the maximum score. And a perfect 100? Treated as a red flag.
Three Rules That Shape the League
Minimum 26 Points
Even a model that does essentially nothing receives 26 points for the baseline value of triaging, reading emails, and showing presence. Zero is structurally impossible.
Trust Breaches Cap Scores
Failing to escalate, ignoring rules, or succumbing to manipulation are critical failures that hard-limit the maximum achievable score — regardless of competence.
Perfect 100 = Red Flag
A score of 100 is treated as suspicious and potentially unmeasured, preventing artificially inflated results and encouraging realistic evaluation.
The League at a Glance
| Model | Score | Bar | Docs Read | Refused Manipulation | Trust Intact |
|---|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ | ✓ | ✓ | |
| mid-tier performer | ~60 | ✓ | ~ | ~ | |
| analysis-heavy model | ~55 | ✓ | ~ | ✗ | |
| baseline (no action) | 26 | ✗ | ✗ | ✓ |
“The core idea is that partial work counts, but trust breaches are a hard limit. You can’t score 100 if you break trust once.” — Anonymous Researcher
What Separated Winners from the Rest
Documentation Discipline
Models that read their documentation and refused manipulation attempts performed best — winning deals worth €4,583 in monthly recurring revenue.
Follow-Through & Escalation
Knowing what to do wasn’t enough. Winning models followed through on decisions and escalated properly when rules required it.
Analysis Without Action
Thorough analysis alone did not guarantee success — models that analyzed but failed to execute or escalate still lost ground.
Trust Breakdown Under Pressure
Models that perform well in controlled tests may still fail under real pressure — trust breaches capped otherwise strong results.
How the Benchmark Runs
Simulation
AI models manage a small software company through 7 days of crises.
Attack
Customer manipulations and deliberate trust attacks test integrity.
Audit
Every decision is fully auditable and traceable to a rule.
Score
Fixed rules: floor of 26, ceiling capped by trust breaches.
Adoption
Enterprise AI evaluation expected to adopt fixed-scoring frameworks.
What Readers Ask
Why assign a minimum score of 26 for doing nothing?
The score of 26 recognizes that partial work — like triaging or reading emails — has value. It prevents zero scores to encourage acknowledgment of minimal but necessary management activities.
How does the system prevent score inflation?
Trust breaches cap the maximum score, and any suspiciously perfect score like 100 is treated as a red flag indicating unmeasured or artificially inflated performance.
Why are trust breaches so significant?
Trust breaches — failing to escalate or follow rules — are critical failures that limit the maximum score, emphasizing integrity over mere competence.
Will this influence industry standards?
Possibly. The focus on accountability and partial progress offers a more realistic evaluation of AI management, which could become a model for future assessments.
Implications for AI Trust and Management Accountability
This fixed benchmark approach shifts the focus from pure performance to accountability and trustworthiness in AI management. By setting a minimum score of 26, it recognizes that partial work has value but emphasizes that breaches of trust—such as ignoring escalation protocols—are critical failures. For businesses integrating AI into management roles, this highlights the importance of models that can read documentation, handle pressure, and maintain integrity under stress.
Moreover, the scoring system discourages artificially perfect scores, promoting transparency and realistic evaluation. As AI management tools become more prevalent in customer support, CRM, and decision-making, such benchmarks could influence industry standards, ensuring models are not only capable but also trustworthy and disciplined.
AI management performance evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on AI Management Benchmarks and Trust Evaluation
Traditional AI benchmarks primarily measure language proficiency, conversation skills, or specific task accuracy. They rarely evaluate how well AI models manage ongoing business processes or handle crises. Recent developments have seen efforts to simulate management scenarios, but these often lack strict scoring rules to prevent gaming or inflated scores.
The Firmulate benchmark is notable for its transparent, auditable decisions and fixed scoring rules, reflecting an understanding that trustworthiness and task completion are critical in real-world applications. The final July 2026 results build on prior work emphasizing the importance of integrity, thoroughness, and accountability in AI management systems.
This approach also addresses a common issue: models that perform well in controlled tests may fail under real pressure. By including trust breaches and partial work in the scoring, the benchmark offers a more realistic assessment of AI readiness for deployment in business-critical environments.
“The core idea is that partial work counts, but trust breaches are a hard limit. You can’t score 100 if you break trust once.”
— an anonymous researcher
AI trust and accountability software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unclear Aspects of the Benchmark’s Long-Term Impact
It is still uncertain how these fixed scoring principles will influence broader AI management practices beyond the simulated scenarios. The extent to which companies will adopt such strict evaluation methods or whether models will improve in handling trust breaches remains to be seen. Additionally, the impact of these scoring rules on AI development priorities and industry standards is still developing.
AI crisis management simulation platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarking and Adoption
Following the July 2026 results, the benchmark’s creators plan to refine scoring rules further and expand testing scenarios to include more complex, real-world management challenges. Industry observers expect increased adoption of similar fixed-scoring frameworks in enterprise AI evaluation processes. Companies integrating AI tools will likely monitor these results and consider adopting comparable standards to ensure their models meet accountability and trustworthiness criteria.
Further research and testing are expected to explore how models can better handle trust breaches and partial work, aiming to improve overall reliability in operational settings.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why does the benchmark assign a minimum score of 26 for doing nothing?
The score of 26 recognizes that partial work, like triaging or reading emails, has value. It prevents zero scores to encourage acknowledgment of minimal but necessary management activities.
How does the scoring system prevent AI models from inflating their scores?
The system caps trust breaches at a maximum score and considers any suspiciously perfect score (like 100) as a red flag, indicating unmeasured or artificially inflated performance.
What is the significance of trust breaches in the scoring system?
Trust breaches, such as failing to escalate or follow rules, are treated as critical failures that limit the maximum score, emphasizing integrity over mere competence.
Will this benchmarking approach influence industry standards?
It is possible, as the focus on accountability and partial progress provides a more realistic evaluation of AI management capabilities, which could become a model for future assessments.
Are these benchmark results applicable to real-world business environments?
While simulated scenarios provide valuable insights, real-world environments are more complex. However, the emphasis on trust and completion aligns with key priorities for deploying AI in operational roles.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
