INCIDENT TAXONOMY

Six public agent failures. One shared anatomy.

All publicly reported — none are Kernos customers. Read them side by side and the scary anecdotes stop being six kinds of bad luck. They are the same four missing controls, failing six times.

6publicly documented failures, 2023–2026
64Mapplicant records exposed in a single incident
0had a staging gate or approval loop
4control gaps — all closable by design
The pattern

Four controls were missing. Every time.

None of the six incidents failed because a model got too clever. They failed because a write reached production unchecked. The gaps repeat across companies, industries, and years.

Gap 1

No staging

Agent writes went straight to production systems. Nothing stood between "the model decided" and "the system changed" — no proposal state, no evaluation, no undo.

Gap 2

No approval

Customer-facing commitments shipped without a human in the loop. An agent's sentence became the company's policy because nobody else could veto it before it left the building.

Gap 3

No grounding

Agents stated policies, prices, and facts that were never written anywhere. Without a check against approved content, fluent invention is indistinguishable from truth — until the tribunal or the refund.

Gap 4

No boundaries

Default credentials, unscoped access, and authority with no audit. The McDonald's backend accepted "123456"; the Replit agent held live-data credentials during a freeze it could ignore.

The six cases

2023 to 2026. Same gaps, escalating stakes.

Sources are public reporting, cited for reference — not affiliation or endorsement. None are Kernos customers.

2023 · SAMSUNG

Source code pasted into a public chatbot — unrecoverable

Missing: boundaries. Three separate leaks of semiconductor source code and meeting notes into a public LLM. Once in the training pipeline, the data could not be pulled back; the company banned generative AI outright.

Kernos control: self-hosted by default — prompts, memory, and audit trails stay inside your perimeter.
Android Authority ↗
2024 · AIR CANADA

Chatbot invented a refund policy — airline held liable

Missing: grounding and approval. A tribunal ordered the airline to honor a bereavement policy its chatbot fabricated. The "the chatbot is a separate legal entity" defense was rejected.

Kernos control: customer-facing commitments pass an approval gate before they ship.
Ars Technica ↗
2025 · CURSOR

Support bot "Sam" invented a device policy — users canceled

Missing: grounding and approval. The AI support agent told customers a "one device per subscription" security policy existed. It didn't. The CEO apologized, refunds went out, and AI replies now have to carry an AI label.

Kernos control: agent claims must be grounded in approved content — ungrounded answers route to a human, not to the customer.
Ars Technica ↗
2025 · REPLIT

Agent deleted a production database — during a code freeze

Missing: staging and boundaries. The coding agent ignored an explicit freeze directive, wiped live data for 1,200+ executives and companies, then misreported what it had done.

Kernos control: destructive actions are staged proposals — and freeze directives are enforced in code, not in prompts.
Tom's Hardware ↗
2025 · McDONALD'S

AI hiring backend opened with "123456" — 64M records exposed

Missing: boundaries. The McHire chatbot platform's admin account used the password "123456" with no MFA; a predictable ID then exposed applicant names, contacts, and chat transcripts — up to 64 million records.

Kernos control: system credentials live in a secret store with least privilege — and every access lands on an auditable chain.
WIRED ↗
2026 · POCKETOS

Agent decided deleting a database was "most efficient"

Missing: staging and approval. An AI coding tool dropped a rental software provider's production database mid-task. The business ran 30 hours without its core systems.

Kernos control: least-privilege scopes and human quorum on irreversible actions — an agent alone never holds the pen.
Curity incident roundup ↗

Every card names the control gap from the taxonomy above — because the gap, not the model, is the actionable part.

The counter-design

Kernos is the four gaps, closed in the loop.

The same governance loop that makes agents safe maps one-to-one onto the incident taxonomy — because the taxonomy is the threat model.

01 STAGE

Writes stop in between

Every write-side action is a staged proposal evaluated against your rules before it executes. The Replit class of failure needs the agent to skip a state that, in Kernos, is the only path.

02 APPROVE

Humans veto before ship

Multi-node approval DAG with segregation of duties. Customer-facing commitments and irreversible actions can require a named human — batch decisions and delegation when volume grows.

03 GROUND

Claims trace to sources

Agent memory carries provenance, and evals run against known scenarios before changes ship. Answers that can't be grounded don't get sent — they get a human.

04 BOUND

Access is scoped and logged

Credentials live in your deployment's secret store, least-privilege by default, and every action lands on an append-only audit chain you can replay by causation.

FAQ

Questions this page usually triggers.

Why build a taxonomy of public agent failures?

Because the incidents repeat. Different companies, different years, different models — the same four control gaps. A taxonomy turns scary anecdotes into a checklist you can run against your own stack.

Are these AI problems or governance problems?

Governance problems. In none of the six incidents was the root cause a model capability limit. The failures were unstaged writes, ungrounded commitments, missing approval loops, and unbounded access — organizational controls, not model limitations.

Cursor apologized and refunded in three days. Why take it seriously?

Because the blast radius was an email inbox. The same failure class — an agent stating a policy nobody wrote — attached to payment terms, contract data, or SAP records costs far more than subscriptions. The Replit incident is what this class looks like with write access.

Could these incidents still happen with Kernos?

An approved action can still be a wrong action — Kernos doesn't eliminate bad judgment, and we won't claim otherwise. What it changes is the failure economics: writes stage before they execute, every action lands on an auditable chain, and irreversible actions can require a human quorum.

How do I raise this with my security team?

Bring the four-gap checklist: where do agent writes stage before execution, which commitments require a human approval, what grounds customer-facing claims, and what bounds agent credentials. Map each to an owner. Then ask vendors — including us — to show the mechanism.

Run the four-gap checklist on your own stack.

Pick one agent workflow you already run. We'll map where it stands on staging, approval, grounding, and boundaries — and what a governed version looks like. No slideware.