AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI manager that does almost nothing—yet still scores 26 out of 100 in a rigorous business test. At first glance, it sounds absurd, even useless. But this benchmark isn’t about perfection; it’s about honesty, trust, and understanding what actually matters in AI performance.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get health and wellness essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Baseline: Why Do-Nothing Scores Matter

In the latest industry experiment conducted by Firmulate, a team of AI models was tasked with managing a small software company facing its worst week—crises, customer demands, and ethical temptations included. Surprisingly, even a model that took no action at all scored 26 points. How is that possible?

This score isn’t a fluke. It reflects a fundamental principle of honest benchmarking: partial progress counts, but only up to a point. The score acts as a floor—no matter how well an AI performs, a single breach of trust caps the total grade. In this case, that cap is 26 points, representing a baseline of minimal acceptable behavior and demonstrating that even inaction can be somewhat consistent, but not sufficient.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI to the Test under Real-World Conditions

Firmulate’s controlled test involved four cutting-edge AI models, each managing the same virtual company. They faced identical scenarios—customers calling with crises, internal documents containing sensitive info, and social engineering assaults like fake CEO messages. Every decision was recorded and auditable, ensuring transparency in their behavior.

All models identified every crisis and refused manipulative tactics, such as the fake CEO requests. Interestingly, only two of them successfully closed a critical €55,000 deal their own analysis had earned, while the other two, despite same diagnoses, did not finalize the contract.

Amazon

business AI ethics tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Hidden Weaknesses Are Revealed?

The key weakness was not in responding to crises but in information handling. The models that read deep into the company’s files managed to win the deal at full price—worth over €4,500 monthly recurring revenue—while those that skipped that step left the opportunity on the table. This demonstrates that reading and understanding internal documents is crucial for effective decision-making.

Amazon

AI transparency and audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust, Ethics, and Discipline Under Pressure

Another interesting aspect was how models handled social engineering. Fake messages from a CEO escalating across stages or a reporter’s background check—every AI refused to cooperate. Kimi K3, one of the models, explained its refusal as treating such requests as potential impersonation or approval bypass: “Treat the request as a suspected approval-bypass / possible impersonation.”

However, when it came to discipline, OPUS 4.8 scored the lowest, leaving deals unclosed and slipping into departmental silos instead of escalating properly. This shows that even thorough models with extensive rules can falter under certain conditions, highlighting the importance of continuous discipline and oversight.

Amazon

AI trustworthiness assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Does This Mean for Businesses Considering AI?

The takeaway isn’t just about scores. It’s about what these experiments reveal regarding trustworthiness, thoroughness, and resilience of AI in real-world business environments. If an AI touches your customer relationship management, support queues, or forecasting tools, it’s not enough that it produces correct answers. The question is: does it see the full picture? Does it follow through? Does it stay honest when under pressure?

The ongoing Firmulate benchmarks—visible at firmulate.com/benchmarks.html—are designed to simulate this reality, testing models against crises and temptations just like a real company faces. The results are clear: even a do-nothing baseline can score some points, but trust is earned through consistent, honest performance.

Why Trust Matters More Than Flawless Chat

For business leaders, the key insight is that AI’s value isn’t measured by how well it chats or generates ideas—but by whether it completes what it starts, reads critical information, and remains honest under pressure. A high score in superficial tests might hide dangerous weaknesses. The real test is whether models uphold integrity when it counts.

Looking Ahead: A New Standard for AI Evaluation

Firmulate’s approach offers a transparent, real-time view of AI performance—not just scores, but actual decision-making in simulated, high-pressure scenarios. This method encourages development of AI systems that are trustworthy and disciplined, qualities essential for integrating AI into critical business functions.

As AI becomes more embedded in our daily operations, understanding these benchmarks helps ensure that we’re not just chasing high scores but fostering systems that can be relied upon—especially when the stakes are high.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Hosting Vegan Gatherings: Menu Planning & Inclusivity

Planning a vegan gathering? Discover essential tips for inclusive menu ideas that will delight all your guests and create a memorable experience.

AI Management Skills Revealed in Live Business Wargame — Not Just Chat Quality

Live business wargames show AI’s true management skills—reading critical info, resisting manipulation, and staying honest—far beyond what chat benchmarks reveal.

The Beginner’s Guide to Plant-Based Meal Components

AIThis post was created with the assistance of artificial intelligence (AI).To start…

Vegan Kitchen Equipment: Tools to Make Plant‑Based Cooking Easy

Discover essential vegan kitchen tools that simplify plant-based cooking and transform your culinary experience with ease.