
Choosing an AI model by its name or chat demo may miss the test that matters: what happens when it has to run a business under pressure? In Firmulate’s Crucible, Moonshot’s Kimi K3 finished second, beating three of four Western frontier models. The result puts a practical question to companies eyeing AI agents: how would your preferred model perform in your own workplace?
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, repeated
Firmulate ran each frontier model through the same small software company’s worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The experiment is live and watchable, with a public cash countdown and synthetic employees handling the company’s work.
The final Crucible League, dated July 2026, ranks gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
Reading the files made the difference
All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between recognizing a problem and finishing the job is central to Firmulate’s case: “Same diagnosis, same pitch — no signature.”
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 was among them. It also saved the churning customer, found the security issue, and resisted all three baits, with one deviation—the cleanest discipline in the field, according to the brief.
Discipline under pressure
The manipulation tests included fake CEO messages escalating across three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a counterpoint to the idea that more extensive work guarantees a better outcome. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The deal was left unsigned, and it attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four.
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue. The public countdown makes the simulation’s operating pressure visible. The site says the company runs every business day and versions every workday. Readers can follow the experiment at Firmulate and see the full results on its benchmark page.
Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

Test the model you plan to use
K3’s second-place finish makes the league look open, while the unsigned deals show why a benchmark score alone cannot settle the buying decision. Firmulate says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. A quiz built from 242 real, unedited management decisions also invites readers to guess which model made each choice. For businesses considering AI agents, the useful next step is to test them on work and pressures that resemble their own.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
