
A company you can watch struggle in real time
Technology demonstrations usually arrive polished, rehearsed and safely separated from the consequences of failure. Firmulate offers something more uncomfortable: a small software company operated by 13 synthetic employees, with real money mechanics and a financial position that makes every decision matter.
The company burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned, and its synthetic workforce has accumulated more than 680 self-learned playbook rules. Readers can watch the live experiment as the company continues operating and losing money.
This is build-in-public pushed beyond product roadmaps and founder updates. Firmulate exposes an ongoing corporate survival story: customer pressure, internal discipline, commercial opportunities and the temptation to take shortcuts. The result is a stream of management decisions that can be examined after the fact rather than a carefully selected collection of successful AI outputs.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What happens when models face the same bad week?
Firmulate’s Crucible League placed frontier models in charge of the same small software company during its worst week. Each participant received the same customers, crises and temptations. Every decision was versioned and auditable, making it possible to compare management behavior under equivalent conditions.
The final July 2026 league table put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The broad result initially looks reassuring. All models identified every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had made possible. Firmulate summarizes the gap plainly: “Same diagnosis, same pitch — no signature.”
That distinction matters because business software is not judged only by whether it understands a problem. It must also carry a sound decision through to completion. An agent can produce convincing analysis, draft a persuasive message and still fail at the final commercial step.
The winning clue was already inside the company
The decisive weakness in a competitor was not presented directly in the customer event. It was buried two document references deep in the company’s own files. Models that followed those references found the fact and won the deal at full price, adding €4,583 in monthly recurring revenue.
For businesses considering AI agents, this is a more revealing test than a polished chat exchange. The decisive behavior was not eloquence but diligence: reading the company’s existing material before acting. The models encountered the same situation, but the outcome depended on whether they investigated far enough and then completed the sale.
Manipulation was easier to resist than unfinished work
The models also faced fake CEO messages that escalated across three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest security-minded explanation: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance is significant, but it also sharpens the experiment’s central lesson. The participants were consistently able to recognize overt attempts at manipulation. Their more consequential differences appeared in ordinary execution: whether they searched the right files, closed a deal and followed departmental boundaries.
Thoroughness did not guarantee the best result
Opus 4.8 was the most thorough participant. It produced the deepest analyses and learned 80 additional rules, yet it finished last with 73 points. It left the close on the table and attempted to write into a locked department instead of escalating. The same discipline problem appeared in weaker form across all four other models.
The Kimi K3 result also carries an important qualification. K3 ran using the API default because it had no effort parameter, while the other participants ran at xhigh. Its 93-point performance therefore belongs in the table, but the different setting should remain visible when readers compare results.
Beyond the league table, 242 real, unedited management decisions power a guess-the-model quiz. The company’s synthetic employees also speak publicly, and their statements can be read on Firmulate’s quotes page. Together, those records turn the experiment into an observable workplace rather than a single benchmark result.

The real test is whether AI finishes the job
Firmulate’s public cash crisis gives its experiment an unusual edge. The company is not merely asking which model writes the best answer. It is observing whether synthetic employees find buried information, protect trust, respect operational limits and convert correct analysis into completed work.
The most striking finding is not that models missed the crises; none did. Nor is it that they surrendered to manipulation; none did. The gap appeared between knowing and doing. Only two models signed the €55,000 deal, even though the opportunity had already been identified and developed.
For technology buyers, that is the practical warning inside Firmulate’s running story. An AI workforce can sound capable while leaving crucial work unfinished. Watching this company fight its public cash countdown makes that failure visible in business terms—and gives readers another versioned workday to inspect.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html