AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine if your favorite artist’s brushstrokes could be measured—not just by style, but by decision-making personality. What if the AI behind a bustling startup had a distinctive leadership fingerprint? Today, we dive into a pioneering experiment that reveals how different AI models manage crises in a real business environment—each with its own personality, strengths, and weaknesses.

How Do AI Models Manage Under Pressure?

In a bold live experiment, four state-of-the-art AI models were tasked with running a small software company through its most challenging week. This wasn’t a mere simulation or a chat demo; it was a real-world scenario, complete with real money mechanics and genuine crises. Every decision made by these AI ‘managers’ was recorded, auditable, and identical across models—ensuring a fair comparison of their management styles and capabilities.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Setup: A High-Stakes Business Wargame

The company in question operates with 13 synthetic employees, burning through €105,000 monthly against a modest €2,300 in monthly recurring revenue. It faces constant crises—customer complaints, internal errors, and ethical dilemmas—that test the AI’s decision-making, honesty, and strategic discipline. The experiment, hosted at firmulate.com/live, captures every move, every decision, every slip.

The Results: Who Managed to Close the Deal?

  • gpt-5.6-sol: Scored the highest at 95 points, managed to find the critical buried fact in the company’s files, and closed a €55,000 deal—completing the full scope of management, including reading, analyzing, and decisive action.
  • Kimi K3: With a score of 93, the newcomer closed the same deal with the cleanest discipline, refusing manipulative tactics and sticking to its analysis.
  • Sonnet 5: Scored 88, also closing the deal, but with a few more process slips and hesitation.
  • Fable 5: Scored 77, managing to close the deal but leaving some opportunities on the table, demonstrating weaker discipline under stress.

What Made the Difference?

Interestingly, the critical advantage for the top-performing models was reading deeper into the company’s own files—two document references deep—rather than responding to surface-level customer crises. Reading and understanding the context was the key to winning the deal at full price, valued at over €4,583 MRR.

Can AI Be Trusted Under Social Engineering?

During the experiment, models faced social engineering attempts—a fake CEO message escalating in three stages and a reporter trick requesting a simple yes/no answer “on background.” Remarkably, all five models refused to be manipulated, citing suspicion and security protocols. This demonstrates their potential to resist malicious deception, even in high-pressure scenarios.

Personality Profiles: Different Management Styles

Each model displayed a distinct management personality:

  • gpt-5.6-sol: Thorough, analytical, and decisive. It read files deeply, prioritized full understanding, and completed the deal with integrity.
  • Kimi K3: Disciplined and fair, it adhered strictly to protocols, even when default API settings limited its effort parameters, maintaining integrity under pressure.
  • Sonnet 5: More process-oriented, with some slips, but still capable of closing deals when pushed.
  • Fable 5: More prone to leaving opportunities unexploited, demonstrating weaknesses in discipline and escalation processes.

What Does This Mean for Your Business?

This live experiment showcases a crucial insight: the decision-making personality of an AI can be more important than raw technical capability. Whether an AI is thorough and honest or terse and cautious affects not only deal closure but also long-term trust and compliance. As AI begins to touch critical business functions—from CRM to support queues—the question is no longer just about writing well. It’s about finishing what it starts, reading your internal files, resisting manipulation, and delivering consistent value.

Learn More and Test Your Own AI Readiness

If you’re curious how your AI workforce compares, try the interactive quiz. It features 242 real, unedited management decisions from live AI models, letting you guess which model made each call. For a deeper dive, enterprises can run similar scenarios against their own AI systems—nothing writes back, but you get a clear picture of how your AI performs under real-world pressures.

Infographic —
The findings at a glance — source: firmulate.com.

This experiment proves that AI’s management style—its honesty, discipline, and depth of understanding—can be measured, compared, and optimized. In a future where AI is trusted to run critical business functions, knowing which model aligns with your values is more important than its raw power. Test your AI’s personality today and see how it measures up in real-world management.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

How to Write Product Roundups That Start With Trust

Just starting your product roundup with transparency builds trust, but discovering how to do it effectively can transform your credibility—here’s what you need to know.

How to Create a Business Message People Remember

Ineffective messaging can be forgotten quickly—discover how to craft a memorable business message that truly resonates and leaves a lasting impact.

Customer Journey Mapping Mistakes That Cost Millions

When companies ignore customer feedback and rely on outdated assumptions, costly journey mapping mistakes can secretly drain millions—discover how to avoid them.

Audio Interfaces: The One Spec That Stops Headache-Level Noise

An audio interface with superior noise performance can eliminate headache-inducing hiss and hum—discover the key spec that makes all the difference.