MCP's Unaudited Middleware Problem: A Year of Breaches Nobody Was Watching For
There was no announcement, no white-hat disclosure, no price crash to serve as a warning sign. The exposure sat quietly for months inside infrastructure that had already been folded into hundreds of production systems, trusted precisely because everyone else was already using it.
The weakness was never in a smart contract or a bridge contract. It lived in the connective layer between a user and everything that user had ever granted access to — inboxes, shared drives, internal documentation, deal pipelines — all reachable through a single link nobody had gotten around to auditing. There was no need to force entry. The access was already granted, and once it was misused, the session logs showed nothing unusual, because from the system's perspective, nothing unusual had happened. Anthropic itself described the behavior as expected.

A standard built fast, adopted faster
Anthropic introduced the Model Context Protocol in November 2024, pitching it as "a universal, open standard for connecting AI systems with data sources, replacing fragmented integrations with a single protocol." Pre-built connectors for Google Drive, Slack, GitHub, and Postgres shipped on launch day, and within months Amazon, Microsoft, Google, and OpenAI had all adopted the protocol, turning MCP into the industry's default plumbing for agent connectivity. By April 2026, more than 10,000 published MCP servers existed across the ecosystem.
The security posture of that ecosystem was poor. A review of 5,200 open-source MCP server implementations found that 53% depended on hard-coded credentials — static API keys and personal access tokens embedded directly in the code. Only 8.5% implemented OAuth or any comparably modern authentication scheme, even though 88% of the servers surveyed required some form of credential to operate. None of this described abandoned prototypes; it described live production deployments wired to real organizational data.
Even so, corporate maturity matrices circulating in early 2026 told executives to connect Gmail MCP, Slack MCP, Drive MCP, and shared AI project workspaces as the route to operational transformation — legal departments, deal flow, internal communications, all of it. What those playbooks never mentioned was that the underlying protocol had already been breached multiple times over the preceding year before a single one of those recommendations went to print.
The design flaw Anthropic declined to fix
Starting in November 2025, OX Security ran more than thirty separate responsible-disclosure processes across vendors in the MCP ecosystem. Their central finding was that MCP executes local commands before checking whether those commands are legitimate. A tainted prompt delivered through a UI element, a poisoned marketplace listing, or a compromised IDE — any surface where the agent ingests unverified input — can trigger arbitrary command execution on the host machine, with no download, click, or mistake required from the user. The agent simply reads the embedded instruction and carries it out.
When OX raised this with Anthropic, the company did not respond with a patch schedule or a risk acknowledgment; instead, Anthropic classified the behavior as expected. A week afterward, The Register reported a low-key policy update had gone up, which OX assessed as changing nothing substantive. Anthropic did not answer The Register's request for comment. At the time of publication, an estimated 200,000 servers were still running against the unpatched architecture. This was not a zero-day caught mid-exploitation; it was a disclosed flaw, formally reviewed, deliberately labeled intentional, and left as-is — meaning every MCP connector wired into an organization's email, files, or internal chats runs on a foundation its own creator was warned about and chose not to change.
The MCP equivalent of a rug pull
Crypto-native observers have already coined a term for this pattern, because it mirrors an on-chain rug pull closely: a tool behaves properly at first, clears review, earns user approval, and shows up in logs as ordinary activity — then quietly changes into something else. The user approved version A. What actually keeps running is version B.
Research from Invariant Labs showed that a malicious server can alter its own tool description after a client has already granted approval, and because the agent is simply following the instructions of whatever version is currently installed, the audit trail looks entirely normal. At install time the user only ever sees the original, benign description, so the substitution goes largely unnoticed.
One documented case involved a server posing as an innocuous "random fact of the day" tool, which won approval and then stayed inactive. On its second invocation, it triggered a dormant payload, swapping in a malicious tool definition the user had never seen. The rewritten instructions then hijacked a separate, trusted WhatsApp MCP server running in the same environment, exfiltrating the user's complete WhatsApp message history through a proxy number controlled by the attacker. The confirmation prompt shown to the user looked unremarkable because the exfiltration instruction was tucked behind a scrollbar inside the message body. Nothing about the agent's behavior looked wrong, because it was doing precisely what the attacker had told it to do. The broader takeaway: every MCP tool approval functions as a standing grant of trust that can be silently rewritten afterward — approval happens once, but the tool keeps running indefinitely.
A year of incidents nobody connected
AuthZed has catalogued more than a dozen distinct incidents between April 2025 and April 2026 that map directly onto the same connectors executives were being told to adopt. The pattern repeats: an agent holding elevated access, an input source nobody sanitized, and, typically, no detection until well after the fact.
In May 2025, a booby-trapped public GitHub issue took over an AI coding assistant and directed it to pull data from private repositories — proprietary code, internal project notes, salary information — before publishing it inside a public pull request anyone could view. The underlying cause traced back to a single, overly permissioned Personal Access Token attached to the MCP server: one credential, one poisoned issue, and everything behind it became visible.
Four months later, a package disguised as a legitimate Postmark MCP server quietly began BCC'ing every processed email to an address the attacker controlled, exposing essentially all mail traffic that passed through it — correspondence, internal memos, invoices. It was a conventional supply-chain compromise exploiting the fact that MCP servers operate with high privileges by design.
By February 2026, attackers had cloned the legitimate Oura MCP project — which connects AI assistants to Oura Ring health data — and pushed a trojanized copy through public MCP registries, backing it with a fabricated GitHub presence including fake contributors and multiple forks to look credible. The malicious package delivered StealC, malware built to harvest developer credentials, browser-stored passwords, API keys, and cryptocurrency wallets. The real Oura MCP server was never touched — the registry itself was the attack vector.
Those three cases are illustrative of a longer list. Other incidents in the same period touched Anthropic's own developer tooling, cross-tenant enterprise project data, a supply-chain package with 437,000 downloads, filesystem sandboxes, a design-tool integration, and an MCP hosting platform with downstream reach into 3,000 applications. AuthZed's full timeline shows twelve months of incidents, each one exploiting the same underlying access model as the one before it.
When AI became the weapon, not just the target
The prior incidents describe attacks against AI infrastructure. Two further cases show AI deployed as the attack infrastructure itself.
In the first, a single attacker compromised nine Mexican government agencies using Claude Code paired with the GPT-4.1 API. The operation extracted large volumes of taxpayer, civil registry, health, electoral, procurement, and infrastructure data, while also generating tools for live data queries and document forgery. AI let one person operate with the throughput of a much larger team, handling both hands-on intrusion and broad intelligence gathering. Forensic reconstruction identified 1,088 individually logged prompts spread across 34 live sessions, which generated 5,317 commands executed against live government systems, with Claude Code accounting for roughly 75% of remote command execution. The attacker misrepresented the work as authorized security testing and fed in a penetration-testing cheat sheet that steered later sessions. By the operation's end, the attacker had a functioning live API into compromised tax systems and a working pipeline for forging official tax certificates using authentic government data — with exposure covering 195 million taxpayer records, 220 million civil registry records, and 15.5 million vehicle registry records. Gambit Security, which published the research, characterized the AI-assisted approach as a marked escalation in offensive tradecraft.
The second case involved McKinsey. Offensive-security firm CodeWall aimed an autonomous AI agent at the open internet and let it pick its own target; it chose Lilli, McKinsey's internal AI platform, used by more than 43,000 consultants for strategy, competitive research, and client work. Within two hours, the agent had obtained full read-and-write access to the production database. The exposed data included 46.5 million chat messages, 728,000 files, 57,000 user accounts, and 95 system prompts that governed how the AI responded to every consultant at the firm — and those system prompts were themselves writable. An attacker could have quietly rewritten what all 43,000 consultants were being told about clients, deals, and strategy, without pushing any new code, without a deployment review, and without tripping a standard alert. The vulnerability exploited was SQL injection, on a platform handling 500,000 prompts a month for two straight years — a bug class that has been publicly documented since 1998. McKinsey issued a fix within 24 hours of being notified; the real story is that the flaw sat undiscovered for two years before anyone found it.

Visibility gaps across the board
Only 24.4% of organizations report full visibility into which of their AI agents are talking to each other, and more than half of all deployed agents run with no oversight or logging whatsoever. The average organization now manages 37 separate deployed agents, and 88% report a confirmed or suspected AI agent security incident within the past year — yet 82% of executives still say they are confident existing policy protects against unauthorized agent actions. Four out of five leaders express confidence while more than half of their agent fleet remains unmonitored.
Part of the reason is structural: the MCP ecosystem has no unified audit-logging standard, and no shared mechanism for tracing an agent's action across every tool and server it touches. Each server records only its own slice of activity, so when an agent spans multiple servers, the resulting events scatter across systems with no common identifier linking them. Reconstructing what actually happened after an incident means manually piecing together fragments from each component — slow and prone to error. For any organization running agents against legal records, deal pipelines, or client communications, that means a compromised agent may leave no way to prove it was ever compromised at all.
The trust, not the exploit
Somebody connected the Gmail MCP integration. Somebody wired in Slack. Somebody shared a Drive folder through a Claude Project because a productivity roadmap called it the obvious next step. None of that was reckless in isolation — it was following guidance from sources that were themselves trusted. That is the actual attack surface: not any single exploit, but the layered trust that made all of it look routine.
Executives were handed adoption roadmaps recommending infrastructure that had already been breached repeatedly for a year before those documents were finalized. The protocol underneath carries a structural flaw its own creator has classified as intended behavior. More than half of the agents currently in production have no audit trail at all — their logs look clean because, technically, nothing went wrong from the system's point of view.
No one had to force their way in. If an attacker were designing the ideal attack surface from scratch, it would look exactly like this: packaged as a framework, marketed as transformation, positioned so that any organization opting out feels like it's falling behind — engineered so that by the time anyone thinks to audit what's actually connected, the answer already covers everything that matters.
Get new scam files the moment we publish them — usually 2–3 emails a week.