
Imagine handing an AI the keys during your busiest service: suppliers wobble, customers complain, a tempting shortcut appears, and one wrong promise could cost you trust. Before testing that on a real restaurant, you would want to see how it handles the pressure. Firmulate is built around that kind of rehearsal for businesses.
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, repeated
Firmulate’s final Crucible League, in July 2026, put frontier models in charge of the same small software company through its worst week. They faced the same customers, crises and temptations. The point was to observe management decisions under pressure, rather than judge a polished answer in a chat window.
The results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league also makes trust a hard boundary: “no amount of good work outweighs a breach of trust.”
Seeing the problem isn’t the same as finishing the job
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding, in Firmulate’s words: “Same diagnosis, same pitch — no signature.”
The deal hinged on a weakness in a competitor’s position buried two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a reminder that acting well depends on finding relevant context, then carrying a decision through.
Pressure also came as fake CEO messages escalating over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness alone did not win
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped: it attempted writes in a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The league is a specific experiment, not a guarantee of how a model will behave in every business.
From watching to trying it on your business
The live Firmulate company has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, alongside a public cash countdown. Its playbooks include more than 680 self-learned rules, and every workday is versioned. Readers can watch the experiment at firmulate.com. A separate quiz turns 242 real, unedited management decisions into a “guess the model” challenge.
For a restaurant group, food supplier or other enterprise, the next step is a pilot using a read-only export of its own business. The exercise tests crisis scenarios against company-specific information and produces a board report with model rankings and weak points in the playbooks. The export is read-only; nothing writes back to real systems.

A model can spot the crisis, reject a con and still fail to close the deal. Firmulate’s experiment makes that gap visible—and offers enterprises a way to rehearse against their own business data before putting AI to work. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
