โ† All AI Failure Cases
Case 002 ๐Ÿ›๏ธ Public Administration / AI Chatbot March 2024

NYC MyCity Chatbot - Illegal Recommendations

Incident status Closed
Dossier status
โœ… LinkedIn Analysis โœ… Infographic available โณ Full case study
EVIDE Case Score Indicative evidentiary assessment (1โ€“5)
Reconstructability
Evidence Survivability
Indep. Verification
Governance Visibility
NYC MyCity Chatbot - Illegal Recommendations - EVIDE Evidentiary Assessment
What happened

New York City launched MyCity, an AI chatbot designed to help businesses navigate city regulations. Independent testing by The Markup revealed the system advised businesses to discriminate against customers, violate labor regulations, and serve unsafe food. The chatbot was eventually retired.

Evidentiary Assessment - 9 questions
What decision failed?
Public-facing regulatory guidance decisions โ€” the system produced legally non-compliant and harmful recommendations to businesses.
What information was available at the time?
Unknown. No public disclosure of training data provenance, retrieval sources, or knowledge cutoff applied to regulatory content.
Which constraints were active?
No externally verifiable record of content safeguards, legal compliance filters, or output review thresholds was published.
Could the failure be reproduced?
Partially. The Markup reproduced specific failure modes through structured prompting - but full reproduction of the original decision context is not possible.
Could an independent reviewer reconstruct the decision months later?
No. The system was decommissioned. No structured evidentiary record of individual responses was anchored externally at generation time.
What evidence survives?
The Markup investigation (published outputs), the City's public statements, and the decommissioning announcement. No structured decision record survives.
What remains unknowable?
How many businesses acted on the illegal recommendations before testing revealed the failures. The full scope of harm is unquantifiable.
Which governance layer failed?
Multiple layers: training provenance governance, output validation, legal compliance review, and human oversight before public deployment.
Which evidentiary properties were missing?
Training provenance, output reconstructability, compliance threshold anchoring, independent audit trail, post-decommission evidence.
Evidentiary properties missing at decision time
โ† All AI Failure Cases