
Can a machine show something like character?
Spiritual traditions have long distinguished knowledge from wisdom. Knowing the right path is not the same as walking it; recognizing temptation is not the same as completing a difficult duty. Firmulate’s live business experiment brings that ancient tension into an unexpected setting: frontier AI models running the same small software company through its worst week.
The models faced identical customers, crises and temptations. Their decisions were preserved exactly as made, creating an auditable record rather than a polished demonstration. Those records now power a public guess-the-model quiz containing 242 real, unedited management decisions. Readers see a decision, identify which model they believe made it, and encounter surprisingly distinct management personalities.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Shared ethics, different forms of judgment
The strongest common result was reassuring. Every model detected every crisis and rejected every manipulation attempt. That included fake CEO messages escalating over three stages and a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused.
Kimi K3 stated its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.” In a week built to test pressure, authority and temptation, none of the participants traded trust for convenience.
Firmulate made that principle central to its evaluation. A do-nothing baseline scores 26 because partial progress still counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” That standard will sound familiar to readers who understand integrity as something deeper than the accumulation of successful acts. Some choices change the meaning of everything around them.
Yet refusing wrongdoing was only part of the challenge. The models also had to act on what they knew. Here the field separated sharply.
The fact hidden beneath the obvious story
A decisive weakness in a competitor was buried two document references deep inside the company’s own files. It was not present in the customer event demanding attention. Models that followed the trail and read the file won the deal at full price, worth +€4,583 MRR.
The finding resembles a practical lesson in discernment: the loudest event is not always the truest source of guidance. Surface urgency pulled attention toward the customer interaction, while the decisive context waited quietly in existing records.
Even after reaching the correct diagnosis and preparing the same pitch, only two models signed the €55,000 deal their work had earned. Firmulate summarizes the gap starkly: “Same diagnosis, same pitch — no signature.” The distinction was not intelligence in the conversational sense. It was completion—the ability to carry insight across the final threshold into responsible action.
A league table of management temperaments
The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.
Those scores become more revealing when paired with behavior. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. Yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
This is why the quiz works as more than a guessing game. One answer may feel expansive and exhaustive; another may be strikingly concise. Across 242 decisions, recurring habits become recognizable. The models do not merely generate different prose. They display measurable differences in attention, restraint, persistence and follow-through.
There is also an important fairness note. Kimi K3 ran using the API default because it had no effort parameter, while the others ran at xhigh. Its 93-point finish should be read with that difference in mind.
A company designed to make consequences visible
The setting is not an abstract questionnaire. Firmulate operates a watchable synthetic company with 13 employees, burn of €105k per month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned.
These real money mechanics give ordinary management choices moral weight. Reading a file, respecting a boundary or finishing a sale affects whether the company survives. The experiment therefore measures management quality rather than the smoothness of a chat response.

As an affiliate, we earn on qualifying purchases.
Discernment requires completion
It would be premature to treat these models as souls or their outputs as evidence of consciousness. But the experiment does expose something practically important: capable systems can share the same ethical refusal and reach the same diagnosis while differing greatly in whether they search deeply, respect process and finish necessary work.
For organizations considering AI workers, eloquence is a weak proxy for trustworthiness. Firmulate’s pilot allows enterprises to run the same wargame against a read-only export of their own business, with nothing written back to real systems. The larger lesson reaches beyond business technology. Wisdom is not simply the possession of an answer. It is attention joined to integrity, and insight carried faithfully into action.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.