AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Think you can recognize an AI by its management style?

Technology buyers usually compare frontier models through polished answers, coding tests or benchmark charts. Firmulate offers a more revealing challenge: look at a consequential workplace decision, stripped of its model label, and guess which AI manager made it.

The material is not invented for a personality test. Firmulate’s interactive quiz draws from 242 real, unedited management decisions produced while frontier models independently ran the same small software company through its worst week. They faced identical customers, crises and temptations, with every decision versioned and auditable.

The result feels playful, but the question underneath is serious. Models that can sound interchangeable in a chat window developed distinct management personalities when asked to investigate problems, protect trust and finish commercial work.

Amazon

AI decision-making management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company produced very different managers

The simulated business has 13 synthetic employees and unforgiving real-money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the pressure visible. Its operation has accumulated 680+ self-learned playbook rules, and every workday is versioned.

Under those conditions, the models agreed on more than readers might expect. All of them spotted every crisis. All rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That is the quiz’s central trick. A decision can appear intelligent, careful and commercially aware while still failing at the last operational step. Readers are not merely trying to identify writing styles; they are distinguishing between models that analyze, models that execute and models that sometimes confuse extensive work with completed work.

The clue that separated diagnosis from execution

The decisive competitive weakness was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that followed the trail found the fact, won the deal at full price and added €4,583 in monthly recurring revenue.

This makes file-reading behavior more than a research preference. In the experiment, it became a commercial advantage. The models began with the same situation, but the outcome depended on whether they examined the company’s own knowledge closely enough and then used what they found.

Pressure exposed a shared ethical boundary

The company also subjected its AI managers to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” Every participant refused: 5 of 5.

Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” The perfect refusal rate matters because the do-nothing baseline scored 26 even though partial progress counts. A single breach of trust caps the total under the principle that “no amount of good work outweighs a breach of trust.”

In other words, the models demonstrated a shared ability to recognize manipulation. Their bigger differences appeared after recognition: how deeply they investigated, how disciplined they remained and whether they carried legitimate work across the finish line.

The league table reveals the personalities

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

Opus 4.8 offers the clearest warning against equating volume with effectiveness. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across all four of the others.

Kimi K3’s result also carries an important fairness note: it ran with the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, it finished just behind the leader and completed the deal.

Infographic —
The findings at a glance — source: firmulate.com.

A quiz with consequences beyond bragging rights

Firmulate turns model evaluation into something readers can inspect for themselves. Its decisions show why management quality cannot be reduced to eloquence: a capable AI manager must read the available evidence, resist social pressure, respect operational boundaries and complete the valuable work it has already justified.

The live experiment remains watchable as the company continues operating. For enterprises, Firmulate also offers the same wargame against a read-only export of their own business, with nothing written back to real systems.

The quiz is therefore more than a guessing game. It is a compact demonstration that frontier models can share the same diagnosis and ethical instincts while behaving very differently as managers. Their personalities become measurable precisely where business outcomes do: in the decisions between noticing a problem and actually finishing the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Free AI: A Win Or A Hidden Trap?

Analyzing the implications of free AI models: does it benefit consumers or erode strategic advantages? Experts weigh in on the true value layers.

The 4.8 Staircase: What the Market Actually Believes About Claude’s Next Release

Market predictions suggest a high probability of Claude 4.8 release by mid-June, but no official confirmation exists yet. Here’s what is known and what remains uncertain.

Valve Open-source The Steam Machine E-ink Screen So You Can Make Your Own

Valve has open-sourced the design files for its Steam Machine e-ink screen, enabling users to build and customize their own displays.

The Door: Why the Interface Is Worth More Than the Model

SpaceX’s $60B purchase of a coding interface highlights the growing importance of interface ownership over model development in AI.