AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Imagine your favorite Bollywood hero in a boardroom: confident, decisive, and unshakeable. Now, what if AI models played that role—making real management decisions in a high-stakes environment? Welcome to the world of AI management wargaming, where the personalities of artificial decision-makers are on full display.

The Live Business Simulator: A Crowdsourced Test of AI Decision-Making

In a groundbreaking live experiment, four leading frontier AI models—including GPT-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were tasked with running a real, money-losing software company through its worst week. The goal? To see if these AI ‘managers’ could handle crises, resist manipulation, and close deals—just like a seasoned CEO facing a turbulent market.

Every decision was transparent and auditable, meaning you could see how each model responded to the same set of challenges: customer issues, internal sabotage attempts, and even staged social engineering tricks. The stakes were high: only two of the four models managed to sign a €55,000 deal their own analysis had earned, despite all four detecting every crisis and refusing to be manipulated.

The Test Results: Different Personalities, Same Tasks

The models scored quite differently on the final leaderboard, revealing distinct ‘personalities’ in management style:

  • GPT-5.6-sol scored a 95—found the buried fact in the company’s files and closed the deal, showing thoroughness and completeness.
  • Kimi K3 scored 93—also closed the deal but with the clearest discipline, refusing all manipulation attempts and treating suspicious requests as potential impersonation.
  • Sonnet 5 scored 88—closed the deal with some process slips, demonstrating solid but less disciplined decision-making.
  • Fable 5 scored 77—also closed the deal but exhibited more slip-ups, often leaving important steps unexecuted.

The experiment also highlighted a hidden weakness: models that delved deeper into internal documents—like GPT-5.6-sol and Kimi K3—secured the big deal at full price, worth over €4,583 in Monthly Recurring Revenue. Conversely, models that overlooked this buried information failed to close the full deal.

Resisting Social Engineering and Ethical Boundaries

In a social engineering test, a staged staged escalation involving staged messages from a fake CEO and an anonymous reporter, all five models refused to proceed with manipulative requests. Kimi K3 explained its stance: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that some AI models are inherently cautious and ethically grounded, even when under pressure.

The Real Business: Money, Mistakes, and Management Under Pressure

The live company in question operates with 13 synthetic employees, handles real money mechanics—burning €105k each month against a monthly revenue of just €2.3k—and runs every workday with versioned rules and open decision logs. It’s an unforgiving environment where AI decision quality directly impacts the bottom line. Yet, despite the rigorous testing, the experiment reveals that even the most thorough AI models can falter under the weight of discipline lapses, as seen with Opus 4.8, which left a critical close on the table.

What This Means for Business and Entertainment Fans Alike

For Bollywood fans used to watching heroes who always close the deal, the experiment is a wake-up call: AI managers need to be trustworthy and disciplined—not just clever. In real-world applications, the questions aren’t about how well AI writes or chats but whether it can finish what it starts, read key documents, stay honest under pressure, and deliver measurable value.

And for companies considering AI to run operations, the takeaway is clear: running a simulation, or “wargaming,” your AI workforce before real deployment can reveal hidden personality quirks that might make or break your bottom line. These experiments aren’t just theoretical—they happen live, every day, at firmulate.com/live.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethical decision-making models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

This Is What Actually Matters in Popcorn Machine for Bollywood Movie Nights

Ineffective popcorn machines can ruin your Bollywood movie nights—discover what truly matters to ensure perfect, hassle-free popcorn every time.

NFT Soundtracks: Why This Film’s Music Is a Crypto Collectible

Pioneering a new era in film music, NFT soundtracks offer fans exclusive ownership and a unique connection to their favorite films—discover the transformative impact they bring.

How Algorithms Picked the Perfect Heroine for This Hit

In an era where algorithms dictate casting choices, discover how data-driven insights shaped the perfect heroine for this blockbuster hit. What secrets lie behind the scenes?

How Short Clips Turn Unknown Bollywood Songs Into Viral Hits

On social media, short clips can rapidly boost unknown Bollywood songs into viral sensations—discover how to leverage this strategy and unlock hidden potential.