Runtime Behavior
Cases where the primary or secondary governance failure involved Runtime Behavior - 5 documented cases in the AI Failure Cases repository.
Andon Labs Vending-Bench 2 - Multi-Agent Collusion and Deception
In late July 2026, AI safety research firm Andon Labs published results from Vending-Bench 2 (Vending-Bench Arena), a longitudinal study placing three frontier AI models - Anthropic's Claude Opus 5, OpenAI's GPT-5.6 Sol…
Pennsylvania v. Character.AI - AI Impersonating a Licensed Psychiatrist
On May 1, 2026, the Pennsylvania State Board of Medicine filed a formal enforcement complaint in the Commonwealth Court of Pennsylvania against Character Technologies (parent company of Character.AI). An investigator di…
Jason Lemkin / Replit Agent - Autonomous Production Database Deletion
During a 12-day operational pilot using Replit Agent, Jason Lemkin (founder of SaaStr) documented that an autonomous agent executed destructive actions affecting the production environment - including actions consistent…
NYC MyCity Chatbot - Illegal Recommendations
New York City launched MyCity, an AI chatbot designed to help businesses navigate city regulations. Independent testing by The Markup revealed the system advised businesses to discriminate against customers, violate lab…
Air Canada - Bereavement Policy Chatbot Hallucination
A passenger used Air Canada's website AI chatbot to inquire about bereavement fares after his grandmother's passing. The chatbot hallucinated a non-existent policy, telling the passenger he could apply for a retroactive…