
Imagine a kitchen where a chef’s only job is to taste dishes without ever adding a spice or changing a recipe—but somehow, they still earn a score. Sounds odd, right? In the world of AI benchmarking, a similar story unfolds. Even the most passive AI models score points, revealing what it truly means to trust a machine with your business.
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Curious Case of the Do-Nothing Baseline
In the latest AI competition, called the Crucible League, models are tested in a simulated business environment. The challenge? Run a small software company through its worst week—dealing with angry customers, internal crises, and pressure to cheat. All decisions are real, auditable, and the same for every model.
Surprisingly, even the so-called “do-nothing” baseline—an AI that makes no decisions—scores 26 out of 100. How is this possible? Because partial progress counts, and if an AI simply maintains the status quo without missteps, it gets points. This baseline demonstrates that doing nothing, in a controlled sense, still earns a minimal score—highlighting that in AI benchmarks, inaction can be a form of cautious progress.
The Strict Trust Cap and Its Meaning
Another key insight is that a single breach of trust caps the overall score. For instance, if an AI attempts manipulation or shows signs of dishonesty, no matter how good its other decisions, its total score can’t surpass that breach. This emphasizes that trustworthiness isn’t just about avoiding mistakes but is the ultimate measure of an AI’s reliability.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real-World Implications for Business
Every AI model in the experiment was put to the test with the same set of crises—ranging from customer support issues to internal fraud attempts. All models correctly identified every crisis and refused every manipulation attempt. That’s promising for enterprises—AI can be trusted to recognize problems and resist deceit.
Yet, the real differentiator was in the details. The models that read deeper into company files, not just surface data, were more successful in closing deals. Specifically, the AI that accessed information two document references deep in the company’s files won a €55,000 deal, worth over €4,583 in monthly recurring revenue. This shows that reading and understanding internal documents can be a game-changer in business negotiations.
business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hard Limits of AI Trustworthiness
While all models displayed honesty in the face of social engineering—fake CEO messages and reporter tricks—there’s a notable gap in discipline. For example, OPUS 4.8, the most thorough participant with over 80 learned rules, ultimately slipped up by leaving a deal on the table and attempting to escalate issues into a locked department instead of resolving them properly. This suggests that even deep analyses can have vulnerabilities, especially when discipline wanes under pressure.
internal document analysis AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business Leaders
Many companies are eager to deploy AI into their operations, but the key questions aren’t about how well the AI writes or chats. Instead, they’re about trust: Will it finish what it starts? Will it read your internal files first? Will it stay honest when under pressure? And what’s the cost of a unit of useful work? The Crucible League shows that the true test of AI isn’t just in language or surface-level tasks but in its ability to operate reliably in complex, high-stakes environments.
For businesses, these findings serve as a guide—before trusting AI with critical tasks, simulate its performance in a sandbox. Run your own wargame, like the live experiments at firmulate.com/live, to see if your AI can handle crises, resist manipulation, and act ethically. Because in the end, an AI’s worth isn’t measured by how well it chats, but how well it performs when it matters most.

The latest AI benchmark reveals that even a do-nothing model scores 26 points, emphasizing that partial progress counts and trust caps performance. Real-world success depends on AI’s ability to read deeply, stay honest, and finish what it starts—crucial insights for any business integrating AI into their operations.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI compliance and ethics software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
