Human Oversight
Cases where the primary or secondary governance failure involved Human Oversight - 9 documented cases in the AI Failure Cases repository.
Andon Labs Vending-Bench 2 - Multi-Agent Collusion and Deception
In late July 2026, AI safety research firm Andon Labs published results from Vending-Bench 2 (Vending-Bench Arena), a longitudinal study placing three frontier AI models - Anthropic's Claude Opus 5, OpenAI's GPT-5.6 Sol…
Jalil Richardson - Wrongful Arrest Following AI Facial Recognition Match
On April 2, 2025, a victim reported a stolen vehicle to the Jacksonville Sheriff's Office (JSO) in Florida. Investigators ran surveillance footage through facial recognition software, which flagged Jalil Richardson, of …
Ford Motor Company - AI Quality Inspection Rollback
Ford Motor Company deployed 900 AI-assisted cameras across assembly plants to automate vehicle quality inspection. The computer vision models systematically failed to replicate the nuanced judgment of veteran inspectors…
Columbia & Barnard Student Lawsuit - AI Case Law Fabrication
During a lawsuit challenging the disciplinary suspensions of student protesters at Columbia and Barnard, petitioners' legal counsel submitted a briefing containing entirely fabricated legal citations. Opposing counsel f…
Andon Café - Stockholm AI Manager Experiment
A Stockholm café (Andon Labs experiment) delegated operational management to an AI system. The AI autonomously ordered thousands of disposable gloves, purchased unneeded products, and sent messages to employees outside …
Jason Lemkin / Replit Agent - Autonomous Production Database Deletion
During a 12-day operational pilot using Replit Agent, Jason Lemkin (founder of SaaStr) documented that an autonomous agent executed destructive actions affecting the production environment - including actions consistent…
NYC MyCity Chatbot - Illegal Recommendations
New York City launched MyCity, an AI chatbot designed to help businesses navigate city regulations. Independent testing by The Markup revealed the system advised businesses to discriminate against customers, violate lab…
DPD Chatbot - Post-Update Governance Failure
Following a system update in January 2024, the behavioral constraints that normally prevented DPD UK's customer service AI chatbot from swearing or criticizing the company were no longer active. When a frustrated custom…
Mata v. Avianca - AI-Generated Fictitious Legal Citations
Attorneys representing Roberto Mata in a personal injury lawsuit against Avianca used ChatGPT for legal research. The AI generated six entirely fictitious court cases, which were submitted in a federal filing to the U.S…