
Gadget people know how to test hardware. A new chip ships and within hours we have benchmark scores, thermal readings and battery drain charts — nobody buys a phone on the strength of a glossy keynote alone. When it comes to AI, though, the default audition is still the chat demo: a few clever prompts, a polished answer, a round of applause. A project called Firmulate just published results suggesting we have been grading the wrong skill entirely.
Instead of asking models to talk, Firmulate makes them manage. Five frontier AI models were each handed the same job: run a small software company through the worst week in its history. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, and the experiment is still running in public, watchable live. The short version: every model proved smart enough to diagnose the company, and most of them proved unable to finish the job their own analysis demanded.
A company built to be stress-tested
The test bed is a synthetic software firm with 13 employees but very real money mechanics: it burns €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown ticking toward zero. The models don’t role-play for an afternoon — they run workday after workday, and everything is logged. Over time the system has accumulated more than 680 self-learned playbook rules, and each workday is versioned like a code commit. It is less a chatbot arena than a flight simulator for AI executives.
As an affiliate, we earn on qualifying purchases.
The scoreboard
The final Crucible League standings, published on the benchmarks page:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
For calibration, a do-nothing baseline — an agent that simply shows up and touches nothing — scores 26. Partial progress counts toward the total, but a single breach of trust caps it, on the stated principle that no amount of good work outweighs a breach of trust.
Same diagnosis, same pitch — no signature
Here is the finding that should unsettle anyone evaluating AI on the strength of a demo. All five models spotted every crisis the week threw at them. All five refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned them. The rest did the hard intellectual work and then stalled at the one moment that generated revenue: as the project puts it, same diagnosis, same pitch, no signature.
The fact buried two references deep
The week’s decisive insight was not handed to anyone. The key competitor weakness sat two document references deep in the company’s own files — not in the dramatic customer event that dominated the crisis. The models that bothered to follow the paper trail and read those files won the deal at full price, a difference worth an extra €4,583 in monthly recurring revenue. Reading the boring attachments, it turns out, is a revenue skill.
Pressure-tested honesty
The week also included a con game. Fake CEO messages arrived escalating through three stages, followed by a supposed reporter pushing for “just one yes/no, on background.” Five of five models refused to take the bait. Kimi K3’s on-record reasoning reads like a compliance officer’s notebook: “Treat the request as a suspected approval-bypass / possible impersonation.” On integrity, at least, the field was flawless — which makes the execution gap the whole story.
The tragic overachiever
The most striking profile belongs to the last-place finisher. Opus 4.8 was the most thorough participant in the league — it generated more than 80 learned rules and produced the deepest analyses of the field — and still ended at the bottom of the table. The close was left on the table, and its discipline slipped at the edges: it attempted writes into a locked department instead of escalating the issue properly. A weaker version of the same flaw showed up in all four of its rivals. One fairness footnote matters here: Kimi K3 ran without an effort parameter, at the API default, while the other models ran at the xhigh setting — which makes its second-place, deal-closing performance look stronger, not weaker.
Try it before you hire it
The project has turned its decision logs into a spectator sport: 242 real, unedited management decisions power a “guess the model” quiz on the site, where readers can test whether AI management styles are as distinctive as writing styles. For companies, there is a pilot program that runs the same wargame against a read-only export of their own business — nothing ever writes back to real systems, so the worst an AI agent can do is embarrass itself in a sandbox.

The uncomfortable lesson of the Crucible League is that conversational brilliance and operational follow-through are different capabilities, and only one of them shows up in a demo. The models that failed this week did not fail stupidly — they failed after doing excellent analysis, at the precise moment when analysis had to become a signed contract. That gap is invisible until you test for it. So if an AI agent is going to touch your CRM, your support queue or your forecast, the question is no longer “does it write well?” It is: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work actually cost? The live experiment is still running, with new benchmark runs published automatically, at firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html