firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Can a machine show something like character?

Spiritual traditions have long distinguished knowledge from wisdom. Knowing the right path is not the same as walking it; recognizing temptation is not the same as completing a difficult duty. Firmulate’s live business experiment brings that ancient tension into an unexpected setting: frontier AI models running the same small software company through its worst week.

The models faced identical customers, crises and temptations. Their decisions were preserved exactly as made, creating an auditable record rather than a polished demonstration. Those records now power a public guess-the-model quiz containing 242 real, unedited management decisions. Readers see a decision, identify which model they believe made it, and encounter surprisingly distinct management personalities.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Shared ethics, different forms of judgment

The strongest common result was reassuring. Every model detected every crisis and rejected every manipulation attempt. That included fake CEO messages escalating over three stages and a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused.

Kimi K3 stated its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.” In a week built to test pressure, authority and temptation, none of the participants traded trust for convenience.

Firmulate made that principle central to its evaluation. A do-nothing baseline scores 26 because partial progress still counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” That standard will sound familiar to readers who understand integrity as something deeper than the accumulation of successful acts. Some choices change the meaning of everything around them.

Yet refusing wrongdoing was only part of the challenge. The models also had to act on what they knew. Here the field separated sharply.

The fact hidden beneath the obvious story

A decisive weakness in a competitor was buried two document references deep inside the company’s own files. It was not present in the customer event demanding attention. Models that followed the trail and read the file won the deal at full price, worth +€4,583 MRR.

The finding resembles a practical lesson in discernment: the loudest event is not always the truest source of guidance. Surface urgency pulled attention toward the customer interaction, while the decisive context waited quietly in existing records.

Even after reaching the correct diagnosis and preparing the same pitch, only two models signed the €55,000 deal their work had earned. Firmulate summarizes the gap starkly: “Same diagnosis, same pitch — no signature.” The distinction was not intelligence in the conversational sense. It was completion—the ability to carry insight across the final threshold into responsible action.

A league table of management temperaments

The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

Those scores become more revealing when paired with behavior. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. Yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

This is why the quiz works as more than a guessing game. One answer may feel expansive and exhaustive; another may be strikingly concise. Across 242 decisions, recurring habits become recognizable. The models do not merely generate different prose. They display measurable differences in attention, restraint, persistence and follow-through.

There is also an important fairness note. Kimi K3 ran using the API default because it had no effort parameter, while the others ran at xhigh. Its 93-point finish should be read with that difference in mind.

A company designed to make consequences visible

The setting is not an abstract questionnaire. Firmulate operates a watchable synthetic company with 13 employees, burn of €105k per month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned.

These real money mechanics give ordinary management choices moral weight. Reading a file, respecting a boundary or finishing a sale affects whether the company survives. The experiment therefore measures management quality rather than the smoothness of a chat response.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI ethics evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Discernment requires completion

It would be premature to treat these models as souls or their outputs as evidence of consciousness. But the experiment does expose something practically important: capable systems can share the same ethical refusal and reach the same diagnosis while differing greatly in whether they search deeply, respect process and finish necessary work.

For organizations considering AI workers, eloquence is a weak proxy for trustworthiness. Firmulate’s pilot allows enterprises to run the same wargame against a read-only export of their own business, with nothing written back to real systems. The larger lesson reaches beyond business technology. Wisdom is not simply the possession of an answer. It is attention joined to integrity, and insight carried faithfully into action.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation games

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Angel Number 333: Creativity and Communication

Lifting your confidence and creative expression, angel number 333 signals divine support—discover how this powerful message can transform your life.

Why Is My Left Ear Ringing Spiritual? Decode the Universe’s Messages

Perplexed by your left ear ringing? Discover the spiritual significance behind this phenomenon and what the universe may be trying to tell you.

Angel Number 1313: Courage and Communication

Lifting your confidence through honest communication, angel number 1313 encourages you to unlock your true potential—discover how to embrace this powerful message now.

555 Is Trying to Tell You Something Big About Your Future

Unlock the hidden message behind 555 and discover how it signals major changes coming your way—continue reading to uncover what your future holds.