AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Imagine handing an AI the keys to a busy animal shelter: adoption appointments, worried owners, supply orders and a sudden outbreak scare all competing for attention. You would want to know how it behaves under pressure before it affects a real animal’s care. Firmulate is testing that question in business—and its results offer a useful warning for anyone considering AI in an animal-facing operation.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pet supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate’s live experiment runs AI models as a small software company, with synthetic employees, customers and real money mechanics. In the final Crucible League, published in July 2026, each frontier model faced the same company and the same worst week. The decisions were versioned and auditable.

The league’s top five were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s standard is deliberately demanding: “no amount of good work outweighs a breach of trust.”

Knowing the right answer wasn’t enough

Every model spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal their own analysis had earned. The gap—“Same diagnosis, same pitch — no signature”—is a reminder that sound advice and follow-through are different capabilities.

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding has a clear parallel for a veterinary practice or animal charity: an agent may recognize a complaint or an urgent need, yet miss the policy, history or contractual detail needed to act well.

Trust held up in a separate test. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of refusal matters anywhere sensitive animal, customer or staff information is involved.

Thoroughness has limits

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the deal on the table and let discipline slip by attempting writes in a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.

There is a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The league is a snapshot of a specific experiment, not a promise about how any model will behave in a shelter, clinic or pet business.

From watching to a pilot

The live company has 13 synthetic employees, burns €105k/month against €2.3k MRR, and publishes a cash countdown. It has accumulated 680+ self-learned playbook rules, with every workday versioned. The experiment is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call.

For animal-care organizations, the useful question is not whether an AI can write a convincing answer. It is whether it can follow your procedures during a crisis, find the relevant information, protect trust and escalate when it lacks authority. Those are questions best explored against your own operating context before an AI touches live work.

Firmulate’s enterprise pilot runs crisis scenarios against a read-only export of your business and produces a board report with model rankings and weak points in your playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The experiment suggests that recognizing a problem is only part of the job: models also need to find the details that matter, complete the right action and respect boundaries. To explore a pilot against your organization’s own data, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Pet-care content is informational — consult your veterinarian for advice about your animal.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Pet Alliance Marks Five Years Since Orlando Shelter Fire And Announces Plans To Honor Cats Lost

Pet Alliance commemorates five years since the Orlando shelter fire and reveals future plans to honor the animals lost in the tragedy.

The Comfort Setup That Helps Dogs Recover After Busy Days

Unlock expert tips to create a cozy recovery space that helps your dog bounce back faster after busy days and discover the secrets inside.

The Pet Buzz | 8/12/26

A nationwide pet food recall was announced today due to contamination concerns, affecting millions of pets. Details are still emerging.

The Washable Cover Detail That Makes Cleaning Easier

Unlock the secret to effortless cleaning with washable covers that resist stains and repel dirt, ensuring your furniture stays pristine longer.