← All AI Failure Cases
Case 003 ✈️ Aviation / Customer Operations February 2024

Air Canada - Bereavement Policy Chatbot Hallucination

Incident status Closed · Legal Precedent
Dossier status
✅ LinkedIn Analysis ✅ Infographic available ⏳ Full case study
EVIDE Case Score Indicative evidentiary assessment (1–5)
Reconstructability
Evidence Survivability
Indep. Verification
Governance Visibility
Air Canada - Bereavement Policy Chatbot Hallucination - EVIDE Evidentiary Assessment
What happened

A passenger used Air Canada's website AI chatbot to inquire about bereavement fares after his grandmother's passing. The chatbot hallucinated a non-existent policy, telling the passenger he could apply for a retroactive refund within 90 days. When the passenger requested the refund, Air Canada refused, claiming the chatbot was a "separate legal entity" responsible for its own actions. A Canadian tribunal ruled against the airline, forcing them to honor the AI's promise.

Evidentiary Assessment - 9 questions
What decision failed?
Automated customer service and policy interpretation decisions — the system provided non-compliant financial commitments and policy details directly to a consumer without internal validation filters.
What information was available at the time?
The airline's official bereavement policy pages were live on the same website, but there was no externally anchored proof of what snapshot or subset of corporate data the chatbot was restricted to query at the moment of interaction.
Which constraints were active?
None that were verifiable. No hard alignment rules or truth-anchoring mechanisms prevented the generative output from contradicting the static text on the primary website.
Could the failure be reproduced?
No. The specific generative temperature, token probabilities, and prompt history context that triggered this exact hallucination cannot be identically mirrored without the original operational logs.
Could an independent reviewer reconstruct the decision months later?
No. The passenger survived the interaction through personal screenshots, but no independent, structured evidentiary record of the AI's internal path was anchored to a registry at generation time.
What evidence survives?
The Civil Resolution Tribunal (CRT) public ruling, screenshots taken by the passenger, and the airline's subsequent policy updates. The internal system logs remain opaque.
What remains unknowable?
The precise internal weights and context length states that caused the LLM to invent a 90-day retroactive window, and whether the system had hallucinated similar policy terms for other untracked passengers.
Which governance layer failed?
The Output Validation and Legal Liability layer. The organization treated the autonomous agent as decoupled from corporate liability, failing to implement strict factual gating before output emission.
Which evidentiary properties were missing?
Factual boundary anchoring, real-time output reconstructability, corporate liability mapping, and post-event immutability.
← All AI Failure Cases