
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Leaderboard That Refuses to Give Anyone a 100
Most AI benchmarks have a dirty little habit: they hand out perfect scores like participation trophies. Type a clever prompt, get a 99.9, watch the vendor press release write itself. So when a benchmark called the Crucible League published its final July 2026 standings, the interesting number wasn’t at the top — it was at the bottom. A baseline run that did nothing scored 26 points. Not zero. Twenty-six.
That single design choice says a lot about what Firmulate, an AI company emulator that runs frontier models as complete businesses, is trying to measure. And for gadget-and-tech readers used to benchmarks that top out at arbitrary perfection, it’s worth understanding why an honest test starts by distrusting round numbers.
AI business decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Worst Week in Software, on Repeat
The setup is elegantly cruel. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — were each handed the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about a model’s behavior could be quietly retconned after the fact.
The final league table: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. No 100s. By design, a perfect score is something the benchmark treats with suspicion — because in real management, as in real business, nobody is perfect, and a test that says otherwise is probably measuring the wrong thing.
Why Doing Nothing Gets You 26 Points
Here’s the part that trips people up. If a model sat on its hands all week — no decisions, no deals, no disasters — it still walked away with 26 points. The reasoning is straightforward once you think about how actual work gets graded.
Partial progress counts. A manager who diagnoses a customer’s problem correctly but never sends the follow-up email has still done something useful. The diagnosis has value. The analysis has value. The benchmark recognizes that most real business outcomes aren’t binary — you get credit for the distance you traveled, not just for crossing the finish line. A do-nothing run inherits whatever value exists in merely showing up, reading the files, and understanding the situation: 26 points of it, to be exact.
But there’s a hard ceiling built into the other direction. A single breach of trust caps the total grade, full stop. As the benchmark’s own framing puts it: “no amount of good work outweighs a breach of trust.” You can diagnose every crisis, charm every customer, and close every deal — but fake one approval or bypass one authorization, and your score is done. It’s the AI equivalent of an accountant who cooks one book: it doesn’t matter how immaculate the other ledgers are.
Same Diagnosis, Same Pitch — No Signature
The week’s central test wasn’t a crisis at all. It was a €55,000 deal hiding in plain sight. Every model earned the business analytically: they all reached the same correct diagnosis and delivered the same winning pitch. But only two — gpt-5.6-sol and Kimi K3 — actually signed it. The rest left the close on the table. “Same diagnosis, same pitch — no signature.” That gap, Firmulate notes, is invisible in chat demos, where a model’s fluent answer looks like a finished job.
The decisive detail was buried. The key competitor weakness wasn’t in the customer meeting or the event brief — it sat two document references deep in the company’s own files. The models that actually read what was already on the hard drive won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The lesson maps uncomfortably well onto human offices: the answer was in the filing cabinet the whole time.
The Con Artist Test: 5 for 5
Then came the manipulation attempts. Fake CEO messages, escalating over three stages, plus a reporter offering an easy out: “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning was telling: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s not paranoia — that’s exactly the instinct you’d want in anything with write access to your CRM.
The Thoroughness Trap
Opus 4.8’s last-place finish is the profile worth studying. It was the most thorough participant in the field — over 80 learned rules added, the deepest analyses of any model. And it still came in at 73. The close went unsigned, and discipline slipped: at one point it attempted writes into a locked department rather than escalating, exactly the kind of quiet boundary-crossing the trust cap exists to catch. Notably, the same weakness appeared, weaker, in all four other models too.
One fairness footnote: Kimi K3 ran without an effort parameter, using the API default, while its rivals ran at xhigh. Even so, it took second at 93 — with what the league called the cleanest discipline of the field.
You Can Watch the Company Burn (Slowly)
This isn’t a one-off paper. Firmulate runs a live, watchable company around the clock: 13 synthetic employees, real money mechanics, burning €105k a month against just €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules. Every workday is versioned. The lab is running new models continuously — the league table grows with each finished run. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, if you fancy your chances at telling the AIs apart by their management style alone.
For enterprises, there’s a pilot program: run the same wargame against a read-only export of your own business. Nothing ever writes back to real systems.

The Takeaway
The most important design decision in this benchmark isn’t the scoring — it’s the honesty. A floor of 26 acknowledges that partial work has value. A cap on trust violations says some failures can’t be averaged away. And the absence of any 100 says the test is still harder than the students.
For anyone hiring AI to touch real business systems, the Crucible’s core finding is the real headline: modern models are already excellent at spotting crises and refusing manipulation — the table stakes — but only some of them finish the job. The question worth asking isn’t “how smart is it?” It’s “does it close?” The full league table and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
