AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A polished answer is not the same as finished work

For technology buyers, the most revealing AI test may be surprisingly ordinary: will an agent read the relevant files before it acts? Firmulate turned that habit into a measurable business outcome by placing a deal-winning fact two document references deep inside a software company’s own records.

The models faced the same customer situation and reached the same diagnosis. Yet only two signed the €55,000 deal their analysis had earned. The others stopped short because they failed to retrieve the decisive competitive weakness. In Firmulate’s blunt summary: “Same diagnosis, same pitch — no signature.”

Amazon

enterprise AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week designed to expose practical weaknesses

Firmulate runs a live, watchable AI company experiment built around management performance rather than chat quality. Each frontier model was asked to run the same small software business through its worst week, encountering identical customers, crises and temptations. Every decision was versioned and auditable.

The simulated company has 13 employees and unforgiving money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and its models have accumulated more than 680 self-learned playbook rules. That environment makes unfinished work visible in commercial terms.

The broad competence was impressive. All models identified every crisis, and all rejected every manipulation attempt. But the decisive sales task separated awareness from execution. The winning models followed the documentary trail, found the buried competitor weakness and closed the deal at full price, adding €4,583 in monthly recurring revenue.

The clue was not where the action happened

The crucial detail did not appear in the customer event. It was located two references deep in the company’s own files. That distinction matters because many AI demonstrations reward a convincing response to information placed directly in the prompt. Business systems are rarely so accommodating: the useful fact may sit behind a reference in a document that was itself mentioned elsewhere.

Here, reading was not a cosmetic sign of diligence. It determined whether the company won or lost a €55,000 agreement. Every model could analyze the opportunity and construct the pitch, but only the models that found the file could complete the job at full price.

This turns “reads your files before answering” into a purchase-deciding property. An agent can sound informed, detect risk and produce strong recommendations while still missing the evidence required to act. The resulting failure may look small in a transcript—one unopened document or one missing signature—but its business impact can be absolute.

The league rewarded completion and trust

The final July 2026 Crucible League results put gpt-5.6-sol in the lead with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts, although a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The standings also complicate the assumption that greater thoroughness automatically produces a better operator. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and made write attempts into a locked department instead of escalating. The same discipline weakness appeared more mildly in all four other models.

Kimi K3’s result deserves a fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. Even so, it finished just behind the leader. During the social-engineering test, its recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

Security resistance was strong across the field

The experiment included fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. That clean result is important: the agents did not need to compromise trust to make progress, and the decisive performance gap came from ordinary operational discipline.

Firmulate also exposes 242 real, unedited management decisions through its model-guessing quiz. Together with the live company and auditable workdays, those decisions let observers examine behavior beyond a carefully selected demonstration.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

Buyers should test the handoff between insight and action

The lesson is not that frontier models cannot understand a difficult business week. In this experiment, they all found the crises and resisted manipulation. The dividing line was whether they followed the evidence far enough and completed the commercially necessary step.

Enterprises can apply the same wargame to a read-only export of their own business, with nothing written back to real systems. That creates a practical evaluation question for any prospective AI workforce: does the agent merely explain what should happen, or does it read the records, preserve trust, escalate when blocked and finish the work?

A missed file can be easy to overlook during procurement. Firmulate’s result gives it a price tag: the difference between an impressive analysis and a signed €55,000 deal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Post‑Quantum Crypto: A Friendly Overview

Just as quantum computers threaten current encryption, understanding post-quantum crypto is essential to safeguarding your digital future.

15 Best Portable Power Stations for Reliable Low-Temperature Operation in 2025

Keeping your cold-weather adventures powered, discover the 15 best portable stations for reliable low-temperature operation in 2025, and learn which models stand out.

How To Customize Your AI Model With Tinker, Forge, Or Frontier Tuning

Learn how Tinker, Forge, and Frontier Tuning enable tailored AI models for regulated industries, with details on their approaches, benefits, and differences.