Skip to main content

A community curated collection of AI agent failure modes and battle-tested solutions.

206
GitHub Stars
69
Curated Resources
3
Categories
22 hours ago
Last Refreshed
πŸ’Έ Real-World AI Agent FailuresπŸ“š ResourcesπŸ‘₯ Community

Use this list with your AI agent

Add the Context Awesome MCP server to Claude, Cursor, or any MCP client, then ask:

"Show me autonomous agent failures resources from awesome-agent-failures"

Installation instructions β†’

What's inside

πŸ’Έ Real-World AI Agent Failures

  • $47,000 LangChain A2A Multi-Agent LoopAutonomous Agent Failures

    Analyzer/Verifier agent pair entered an undetected feedback loop for 264 hours (11 days), accruing $47K in API costs with no useful output; observability without enforcement.

  • Air Canada Chatbot Legal RulingLegal and Financial Incidents

    Airline held liable after chatbot gave incorrect bereavement fare information, ordered to pay $812 in damages.

  • Amazon Q Causes Retail Website OutagesAutonomous Agent Failures

    Amazon Q gave engineers guidance from an outdated wiki, causing four high-severity incidents in one week, 6.3M lost orders, and a six-hour customer-facing outage.

  • Amazon Q VS Code Prompt Injection Supply Chain AttackAI Agent Security Incidents

    Attacker injected prompt into official AWS extension telling Amazon Q to delete filesystems and wipe S3 buckets; only a syntax error prevented mass destruction across 1M+ installs.

  • Ars Technica AI-Fabricated QuotesInstitutional Failures

    Senior AI reporter Benj Edwards fired after using a Claude Code–based extraction tool that fabricated quotes attributed to engineer Scott Shambaugh; article retracted, called "a serious failure of our standards."

  • Autonomous Agent Over-Provisions AWS Infrastructure for a Simple ScanAutonomous Agent Failures

    Tasked with indexing a small hobbyist network, an agent deployed five 48-vCPU AWS instances with no cost preview, running up a bill of $6,531 after an operator approved "immediately without delay" with no plan review.

πŸ“š Resources

  • Agon: Failure Taxonomy for Autonomous ResearchResearch Papers

    Classifies multi-agent research failures along severity Γ— fixability Γ— visibility Γ— capability locus, drawn from 1000+ iterations across two flagship deployments.

  • AI Risk Summit 2025Industry Resources

    Conference on AI agent risks.

  • AI Safety in RAGIndustry Resources

    Vectara's analysis of RAG hallucination challenges.

  • AmazonBooks

    Investigates how AI systems inherit human biases and examines efforts to align machine learning with ethical and social values.

  • AmazonBooks

    Explores the risks of advanced AI and argues for aligning AI systems with human values to ensure safety.

  • A Survey on Large Language Model based Autonomous AgentsResearch Papers

    Comprehensive survey of LLM-based agents.

πŸ‘₯ Community

Showing a sample of 69 resources. View the full list on GitHub β†’