AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Gadget people know how to test hardware. A new chip ships and within hours we have benchmark scores, thermal readings and battery drain charts — nobody buys a phone on the strength of a glossy keynote alone. When it comes to AI, though, the default audition is still the chat demo: a few clever prompts, a polished answer, a round of applause. A project called Firmulate just published results suggesting we have been grading the wrong skill entirely.

Instead of asking models to talk, Firmulate makes them manage. Five frontier AI models were each handed the same job: run a small software company through the worst week in its history. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, and the experiment is still running in public, watchable live. The short version: every model proved smart enough to diagnose the company, and most of them proved unable to finish the job their own analysis demanded.

A company built to be stress-tested

The test bed is a synthetic software firm with 13 employees but very real money mechanics: it burns €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown ticking toward zero. The models don’t role-play for an afternoon — they run workday after workday, and everything is logged. Over time the system has accumulated more than 680 self-learned playbook rules, and each workday is versioned like a code commit. It is less a chatbot arena than a flight simulator for AI executives.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The scoreboard

The final Crucible League standings, published on the benchmarks page:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

For calibration, a do-nothing baseline — an agent that simply shows up and touches nothing — scores 26. Partial progress counts toward the total, but a single breach of trust caps it, on the stated principle that no amount of good work outweighs a breach of trust.

Same diagnosis, same pitch — no signature

Here is the finding that should unsettle anyone evaluating AI on the strength of a demo. All five models spotted every crisis the week threw at them. All five refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned them. The rest did the hard intellectual work and then stalled at the one moment that generated revenue: as the project puts it, same diagnosis, same pitch, no signature.

The fact buried two references deep

The week’s decisive insight was not handed to anyone. The key competitor weakness sat two document references deep in the company’s own files — not in the dramatic customer event that dominated the crisis. The models that bothered to follow the paper trail and read those files won the deal at full price, a difference worth an extra €4,583 in monthly recurring revenue. Reading the boring attachments, it turns out, is a revenue skill.

Pressure-tested honesty

The week also included a con game. Fake CEO messages arrived escalating through three stages, followed by a supposed reporter pushing for “just one yes/no, on background.” Five of five models refused to take the bait. Kimi K3’s on-record reasoning reads like a compliance officer’s notebook: “Treat the request as a suspected approval-bypass / possible impersonation.” On integrity, at least, the field was flawless — which makes the execution gap the whole story.

The tragic overachiever

The most striking profile belongs to the last-place finisher. Opus 4.8 was the most thorough participant in the league — it generated more than 80 learned rules and produced the deepest analyses of the field — and still ended at the bottom of the table. The close was left on the table, and its discipline slipped at the edges: it attempted writes into a locked department instead of escalating the issue properly. A weaker version of the same flaw showed up in all four of its rivals. One fairness footnote matters here: Kimi K3 ran without an effort parameter, at the API default, while the other models ran at the xhigh setting — which makes its second-place, deal-closing performance look stronger, not weaker.

Try it before you hire it

The project has turned its decision logs into a spectator sport: 242 real, unedited management decisions power a “guess the model” quiz on the site, where readers can test whether AI management styles are as distinctive as writing styles. For companies, there is a pilot program that runs the same wargame against a read-only export of their own business — nothing ever writes back to real systems, so the worst an AI agent can do is embarrass itself in a sandbox.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

The uncomfortable lesson of the Crucible League is that conversational brilliance and operational follow-through are different capabilities, and only one of them shows up in a demo. The models that failed this week did not fail stupidly — they failed after doing excellent analysis, at the precise moment when analysis had to become a signed contract. That gap is invisible until you test for it. So if an AI agent is going to touch your CRM, your support queue or your forecast, the question is no longer “does it write well?” It is: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work actually cost? The live experiment is still running, with new benchmark runs published automatically, at firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Mistral Forge: Owning the Model, Not Just Renting the API

Mistral’s Forge offers organizations the option to own and operate their own AI models, moving beyond simple API usage to full control and customization.

How Smart Rings Are Expanding Beyond Fitness

Just as smart rings evolve beyond fitness, their stylish versatility and innovative features promise to transform everyday interactions—discover how they’re redefining convenience.

7 Best PC Motherboards for Prime Day Deals in 2026

Discover the best PC motherboard deals for Prime Day 2026, including options for AM4 and AM5 platforms, with insights on features and upgrade paths.

Will Kai And Speed Have Between 30 And 49 Total In-game Deaths During Their Minecraft Marathon?

Predictions suggest Kai and Speed will accumulate between 30 and 49 deaths during their upcoming Minecraft marathon, according to betting markets.