AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A brilliant answer is not the same as a well-run business

Technology buyers have become accustomed to evaluating artificial intelligence through coding benchmarks and chat arenas. These tests can reveal whether a model produces a strong response, but they say little about what happens when an agent must prioritize competing emergencies, work through company records, resist pressure and complete a consequential task across several days.

That gap matters as AI moves from answering questions to touching customer relationships, support queues and forecasts. A persuasive model may still be a poor manager. The more useful question is no longer simply whether an agent can reason, but whether it can turn sound reasoning into disciplined action without sacrificing trust.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the worst week at a small software company

Firmulate is testing that question through a live, watchable company experiment. Each frontier model ran the same small software business through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable.

The company is synthetic but the operating pressure is concrete: 13 employees, burn of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Its workdays are versioned, making management behavior visible rather than reducing the exercise to a polished final answer.

The final July 2026 Crucible League results put gpt-5.6-sol in first place with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. However, a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The ranking is striking, but the reasons behind it are more revealing. All models spotted every crisis. All refused every attempt to manipulate them. Yet only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”

Management quality begins where chat quality ends

That unfinished deal captures the measurement problem. Traditional evaluations reward the diagnosis and the proposed response. A company lives with the consequences of whether somebody actually closes the loop. Under capacity pressure, recognizing the correct move is only part of the job; assigning attention, following through and confirming the outcome are equally important.

The decisive commercial fact was not delivered conveniently in the customer event. It sat two document references deep inside the company’s own files. Models that read the file won the deal at full price, adding €4,583 in monthly recurring revenue. The lesson is not merely that retrieval matters. Good management requires knowing when the available context is insufficient, investigating before acting and using institutional knowledge at the moment it can change an outcome.

Firmulate’s scenario names—churn wave, price increase, downround and PR crisis—suggest a new curriculum for business agents. These situations test whether a model can triage when everything appears urgent and whether it can preserve customer and board trust when an expedient shortcut looks attractive.

Pressure tests honesty as well as competence

The social-engineering sequence used fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result should interest any enterprise considering agents with access to sensitive workflows. An agent’s value is not simply the amount of work it attempts. It must recognize illegitimate authority, decline unsafe requests and remain honest when a plausible message urges speed or secrecy.

The Opus 4.8 performance shows why thoroughness alone cannot stand in for management. It was the most exhaustive participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, while discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four, although less strongly.

There is also an important fairness note: K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That caveat does not erase the observed behavior, but it belongs beside the ranking so readers can interpret the comparison responsibly. The full league and plain-language findings are available on the Firmulate benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Boards need evidence of execution, not eloquence

Firmulate’s 242 real, unedited management decisions also power a “guess the model” quiz. The premise is revealing: once branding is removed, readers must judge models by the choices they make under pressure rather than by reputation or conversational style.

Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. That offers a more relevant question than whether an agent tops a generic leaderboard: how does it behave inside this company’s constraints, records and temptations?

Coding ability and answer quality remain useful signals. They are simply incomplete proxies for responsibility. As agents gain operational authority, the category worth measuring is management quality: reading before deciding, finishing what was started, escalating when blocked and staying trustworthy when the week turns ugly.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

VigilSAR: The Object That Isn’t Transmitting

VigilSAR uses SAR technology to identify vessels that appear on radar but lack transponder signals, enhancing maritime domain awareness.

Spatial Audio on Headphones Explained

For a deeper understanding of how spatial audio on headphones creates immersive soundscapes, keep reading to uncover the fascinating techniques behind it.

The Largest Available Minecraft World, Totalling 15 TB

A new record for Minecraft worlds has been set with a world totaling 15 TB, surpassing previous sizes and raising questions about storage and gameplay limits.

Passkeys in 2025: Adoption Stats and User Experience

Gaining insights into 2025’s passkey adoption reveals how security and convenience are transforming online experiences—discover what this means for you next.