
Dogs and AI: What Can a Coding Leader Teach Us About Business Resilience?
Just like a loyal pet that responds under pressure, a business’s success depends on more than just its shiny exterior or quick responses. It’s about how well it manages crises, stays honest, and follows through — especially when stakes are high. Recent experiments with artificial intelligence models in running a simulated company reveal surprising lessons that go beyond chatty demos and into the heart of management quality.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: A Company Under Pressure
Imagine a real small software firm facing its worst week: angry customers, internal crises, tempting shortcuts. Four advanced AI models were tasked with running this company, each facing identical challenges, with the same customers and crises. The goal? To see if the AI could navigate through the chaos, maintain honesty, and close deals at full value.
The results? All four AI models detected every crisis and refused to be manipulated — a promising sign. But only two managed to sign the €55,000 deal their own analysis had earned. The others faltered, failing to finish the job despite sound diagnoses. The key weakness? Hidden in the company’s files, not the visible customer interactions. Models that read and interpret these internal documents secured the deal at full price, adding €4,583 MRR.
What Chat Demos Miss
This experiment exposes a critical gap: traditional benchmarks, like chat-based performance, don’t measure management qualities such as follow-through, honesty, or resilience under pressure. A chatbot that answers well in demos might still collapse when faced with real-world crises and temptations. The AI’s ability to read internal documents and uncover hidden facts proved decisive — a skill often invisible in standard evaluations.
Managing Under Fire
When the AI was tested against social engineering, like fake CEO messages escalating over stages, all models refused to comply. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that true management involves skepticism, verification, and ethical judgment — qualities beyond simple answer quality.
internal document analysis AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human-AI Test Bed
The live company run by AI involves 13 synthetic employees, handling real money mechanics — burning €105k per month against €2.3k MRR. Every workday, the AI models make decisions, learn, and adapt, with over 680 self-learned rules. It’s a transparent, ongoing experiment that anyone can watch at firmulate.com/live.
Interestingly, performance varied across models. Opus 4.8, the most thorough participant with over 80 learned rules, left deals on the table and slipped discipline — showing that thoroughness alone isn’t enough if focus and follow-through waver. The leaderboard? GPT-5.6-sol scored 95, uncovering the critical hidden fact and closing the deal; Kimi K3 followed closely with 93, demonstrating the importance of fairness and discipline.
Why Management Quality Matters
This experiment points to a vital insight: success in real-world AI management isn’t just about answering correctly. It’s about reading deeply, resisting manipulation, making disciplined decisions, and completing tasks regardless of temptations. These qualities are rarely reflected in simple chat benchmarks but are essential for AI to be truly trustworthy and useful.
business crisis management AI solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Your Business
If AI agents will someday handle your CRM, support, or forecasting, the question isn’t just whether they sound convincing. It’s whether they can finish what they start, read and interpret critical internal data, and stay honest under pressure. And crucially, what does a unit of useful work cost?
Firmulate’s ongoing live experiment demonstrates that measuring management skills directly — in real, complex scenarios — reveals gaps invisible to chat demos. It’s a call for better benchmarks, better testing, and smarter deployment of AI in business-critical roles.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethical decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.