AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Choosing an AI model by its name or chat demo may miss the test that matters: what happens when it has to run a business under pressure? In Firmulate’s Crucible, Moonshot’s Kimi K3 finished second, beating three of four Western frontier models. The result puts a practical question to companies eyeing AI agents: how would your preferred model perform in your own workplace?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

Firmulate ran each frontier model through the same small software company’s worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The experiment is live and watchable, with a public cash countdown and synthetic employees handling the company’s work.

The final Crucible League, dated July 2026, ranks gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI model testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading the files made the difference

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between recognizing a problem and finishing the job is central to Firmulate’s case: “Same diagnosis, same pitch — no signature.”

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 was among them. It also saved the churning customer, found the security issue, and resisted all three baits, with one deviation—the cleanest discipline in the field, according to the brief.

Discipline under pressure

The manipulation tests included fake CEO messages escalating across three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a counterpoint to the idea that more extensive work guarantees a better outcome. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The deal was left unsigned, and it attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four.

Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue. The public countdown makes the simulation’s operating pressure visible. The site says the company runs every business day and versions every workday. Readers can follow the experiment at Firmulate and see the full results on its benchmark page.

Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the model you plan to use

K3’s second-place finish makes the league look open, while the unsigned deals show why a benchmark score alone cannot settle the buying decision. Firmulate says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. A quiz built from 242 real, unedited management decisions also invites readers to guess which model made each choice. For businesses considering AI agents, the useful next step is to test them on work and pressures that resemble their own.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Show HN: OpenTIE And OpenXWA, Modern Ports Of Tie Fighter And X-Wing Alliance

OpenTIE and OpenXWA bring updated, playable versions of Tie Fighter and X-Wing Alliance to modern systems, enhancing accessibility for fans and new players.

9 Best Computers, Tablets & Components for Everyday Computing in 2026

Discover the best computers, tablets, and components for everyday use in 2026, based on expert rankings and current market offerings.

The Earnings Call Gap: What Q1 2026 Just Told Us About AI ROI

Analysis of Q1 2026 earnings shows a widening gap between AI investment claims and measurable ROI, affecting stock performance and investor confidence.

Will Kai And Speed Beat The Minecraft Challenge By August 12?

Kai and Speed aim to beat the Minecraft challenge by August 12, with betting markets indicating high confidence. The outcome remains uncertain.