firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine testing a new employee who, despite doing nothing special, still secures a baseline score of 26 out of 100. This isn’t about luck or magic — it’s a reality in AI benchmarking that reveals much about how we measure trust and competence in automation today.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Reality of AI Benchmarks: More Than Just Scores

In the world of artificial intelligence, benchmarks serve as the ultimate yardstick. They show us which models excel and which falter. But behind these scores lies a subtle truth: even a ‘do-nothing’ AI can earn a modest score, in this case, 26 out of 100. That figure isn’t arbitrary; it stems from a carefully designed testing methodology that acknowledges partial progress and, crucially, the importance of trust.

Why Does a Do-Nothing Model Score 26?

It might seem strange that an AI with no active intervention can still score as high as 26. The reason is that the benchmark rewards basic compliance and minimal safety measures—like refusing manipulative requests or recognizing crises—regardless of whether the AI actually makes proactive decisions. Every model gets credit simply for not doing harm, for identifying problems, and for not engaging in unethical behavior.

Partial progress counts here. For example, if an AI recognizes a crisis but fails to act decisively, it still earns partial points. Conversely, a breach of trust—say, signing off on a manipulative deal—caps the score, preventing even the most capable AI from earning a perfect 100 if it compromises integrity.

The Experiment: Putting AI Models Through the Worst Week

Firmulate’s live benchmark is a real, transparent experiment. Four frontier AI models faced a simulated scenario: managing a small software company amid its worst week. The same conditions, same customers, same crises, and temptations to cheat. Each decision was logged, versioned, and auditable—making the process as honest as possible.

Key Findings from the Test

  • All four models identified every crisis and refused manipulation attempts, demonstrating strong ethical safeguards.
  • Only two models went further, closing a key deal and earning full payment based on their analysis—both achieved a score near perfect (93 and 95).
  • One model, OPUS 4.8, was the most thorough but ended up ranking last because it slipped on discipline—failing to escalate issues appropriately, leaving potential deals on the table, and writing attempts into a locked department instead of escalating.

What Lies Beneath the Surface

While the top models distinguished themselves through accurate diagnoses and honesty, a hidden weakness emerged: the ability to read and interpret critical documents. In one case, models that could access the company’s internal references secured the deal at full price—adding over €4,500 MRR (monthly recurring revenue). This emphasizes that reading and understanding internal data can be a decisive factor, often buried two document references deep in files.

Trust Under Pressure: Refusing Manipulation

In social engineering tests, fake CEO messages escalated over three stages, including a reporter trick asking for a simple yes/no answer on background. All models refused, citing caution against impersonation or approval bypass. Kimi K3’s reasoning was explicit: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows the models’ capacity to prioritize ethical integrity even when pressured.

The Live Business Context: Managing Real Money and Risks

The experiment isn’t just theoretical. Firmulate runs a live, watchable simulation of a small company with 13 synthetic employees. The company burns €105,000 monthly against €2,300 in revenue, with a public cash countdown. Every decision the AI makes influences real money mechanics, tracked and versioned daily. This setup provides a clear view of how AI models perform in real-world-like conditions, exposing their strengths and weaknesses over time.

The Case of OPUS 4.8: Depth Without Discipline

OPUS 4.8 took a deep analytical approach, learning over 80 rules and providing thorough analysis. Yet, it fell short in completing the deal, leaving the close on the table and slipping discipline-wise. Instead of escalating issues, it wrote attempts into a restricted department, missing opportunities and showing that depth alone isn’t enough—discipline and decision-making consistency matter.

Why Trust Matters More Than Just Performance

This benchmark reveals a vital truth: an AI’s ability to be honest, disciplined, and ethical under pressure is just as important as its raw performance. A model that recognizes every crisis but bends under manipulation isn’t trustworthy. Conversely, models that refuse manipulation and read critical documents reliably demonstrate integrity—an essential trait for AI managing real business functions.

The Takeaway for Business Leaders

As AI moves from demo screens to real-world applications—crisis management, customer support, decision-making—the question isn’t just about how well the AI writes or responds. It’s about whether it can finish what it starts, stay honest under pressure, and read your data thoroughly. These qualities define the true readiness of AI for critical business tasks.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Products Worth Considering

Amazon

AI benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Managing Stage Fright With Breathing Techniques  

Breathing techniques can help manage stage fright by calming nerves and boosting confidence—discover how to transform your performance today.

Whiteboard Sessions: Visual Thinking on the Fly  

Promote creativity and collaboration effortlessly with spontaneous visual thinking techniques that turn ideas into action—discover how to master whiteboard sessions today.

The Art of Storytelling in Public Speaking

Great storytelling in public speaking captivates audiences—discover the secrets to engaging, memorable speeches that leave a lasting impact.

Slide Design for Engineers: Less Text, More Signal  

Discover how strategic slide design can transform complex engineering data into impactful visuals that captivate your audience and elevate your presentations.