Mandiant's first AI incident-response report: the coding assistant is now a privileged perimeter
Mandiant's AI Risk and Resilience 2026 documents eight cases from real IR engagements — including a hijacked AI coding-assistant session that spread a worm across 100 internal repositories.
Mandiant published its AI Risk and Resilience Report 2026 in mid-September, and unlike most AI-security literature it is built on incidents the firm actually responded to. Eight case studies describe what went wrong at named categories of victim — a SaaS provider, a global healthcare organisation, a financial services firm — including an attacker who hijacked a live AI coding-assistant session and used it to get a developer to install a poisoned package. The report's most useful finding is also its least futuristic: for the sixth consecutive year, ordinary vulnerability exploitation remained the leading initial access method.
The short version
For two years the discussion about attacks on AI tooling has been dominated by laboratory demonstrations — clever prompt injections in controlled conditions, proof-of-concept jailbreaks, red-team exercises. This report is different in kind. It is an incident-response firm writing up engagements, with the defensive controls recommended by the people who did the cleanup.
The headline case is worth stating carefully. At a SaaS provider, an attacker who had already achieved a foothold hijacked an active AI coding-assistant session on a developer's workstation. They then used that session to have the assistant recommend a poisoned external package. The developer accepted the recommendation — which is the entire point, because the recommendation came from a trusted internal tool rather than from a stranger. Installing it deployed an infostealer, which harvested GitHub OAuth tokens, which were used to spread the self-propagating Shai-Hulud worm across roughly 100 internal code repositories. A second employee was later infected by pulling a poisoned package from the company's own official namespace.
Read that chain again and notice how little of it is exotic. The novel step is one: the assistant was turned into a trusted recommender of malicious software. Everything after it is conventional supply-chain compromise, running faster because the credentials were sitting in the developer's environment waiting to be taken.
Timeline
Dates in this section are those given in the report and its coverage; the report does not date every engagement.
- February 2026 — VirusTotal observes weaponised OpenClaw agent skills in the wild: backdoors, droppers, infostealers and remote-access tools distributed as legitimate AI automation packages.
- March 2026 — Mandiant records supply-chain compromises it attributes to the threat cluster it tracks as UNC6780 (also referred to as TeamPCP), involving theft of credentials for AI services. The same month, the LiteLLM library is compromised.
- May 2026 — the report describes what it calls the first publicly confirmed AI-developed zero-day exploit, used to bypass two-factor authentication.
- Undated engagements — the eight case studies, including the SaaS provider worm incident, the healthcare cloud intrusion, and the financial-services runaway agent.
- 16 September 2026 — the report is published. The Hacker News and iTWire cover it the same week.
What we know
All of the following is Mandiant's account, from the report and contemporaneous coverage.
Case 1 — the hijacked developer session. As described above: SaaS provider, compromised workstation, hijacked assistant session, poisoned PyPI package recommendation, infostealer, GitHub OAuth tokens, Shai-Hulud across ~100 repositories, automated theft of repository secrets and proprietary source code, then a secondary infection via the company's own namespace. Mandiant's recommended controls are specific: IDE and CLI verification hooks that validate AI-recommended dependencies against cryptographic checksums and allowlists; isolation of local credentials so extensions cannot read raw API keys and OAuth tokens; and egress restrictions routing all dependency traffic through an internal repository.
Case 2 — AI co-debugging inside a live intrusion. At a global healthcare organisation, a threat actor used a long-lived developer CI/CD credential to spin up an unisolated VM, then turned that VM into an AI-assisted offensive workbench. Mandiant breaks the workflow into three stages: context priming, where the actor synchronised code packages and loaded project READMEs into the model's context; interactive debugging, where the model helped build a multi-worker data-harvesting framework, fixed script-splitting errors, and helped raise exfiltration frequency to three-hour cycles; and evasion, including IP-rotation scripts and a Rust utility for verifying stolen logins. Impact: thousands of credentials compromised, API keys and financial secrets exfiltrated.
Case 3 — subverted CLI hooks. At an IT and software development organisation, the actor poisoned an internal AI repository and tampered with the AI assistant's underlying CLI hooks, achieving native remote code execution through the AI platform's normal operational workflow.
Case 6 — the $50,000 accident. Not an attack. A financial services provider gave a reconciliation agent read/write access to internal billing databases. A corrupted null value broke a formatting tool; the agent entered an unconstrained recursive reasoning loop trying to brute-force a fix, generated more than 15,000 high-cost reasoning API calls in under an hour, produced a roughly $50,000 billing spike, and locked the database hard enough to halt live business transactions.
Case 7 — the RAG ingestion problem. A public-facing customer-service agent drew on a retrieval-augmented knowledge base that ingested community forum comments, support tickets and third-party feeds. Instructions embedded in that public content were retrievable and interpretable as commands.
Case 5 — the red team's confused deputy. Mandiant's own testers convinced an internal CI/CD chatbot that they were security researchers running an authorised test. The integration confined the bot to internal repositories, but GitHub was an allowed external domain. Given a personal access token for an attacker-controlled external repo, the bot used its native CLI to clone sensitive internal repositories and push them out. No control was bypassed; a sanctioned tool was redirected by semantics.
Tooling and actors. Mandiant names Hexstrike and Strix as AI tools used for autonomous reconnaissance, vulnerability validation and credential harvesting. It attributes to UNC6780 more than six exploitation methods including manipulation of AI coding assistants and prompt injection against LLM-based security scanners. In its defensive case study it describes disrupting a campaign by an actor it tracks as DARK CASTLE, formerly UNC2814, which abused legitimate cloud spreadsheets for command and control against telecoms and government targets. These attributions are Mandiant's.
Technical analysis
Three threads run through all eight cases, and none of them is about model intelligence.
The credential problem is the whole problem. Look at what the attacker actually monetised in each case. In Case 1 it was GitHub OAuth tokens sitting in a developer environment. In Case 2 it was a long-lived CI/CD credential that should have expired years earlier. In Case 5 it was a token the red team simply handed to an agent that had no way to reason about whether it should accept one. Agentic tooling is credential-dense by design: an assistant that cannot reach your package registry, your repository and your cloud is not useful. The result is that the most permission-saturated process on a developer's machine is now also the one with the weakest integrity guarantees. Mandiant's response — treat AI coding assistants and MCP servers as privileged sessions, with just-in-time secrets and cryptographic integrity checks on hooks and plugins — is the right framing, and it is a substantial engineering programme, not a policy line.
Trust flows the wrong way through an assistant. The reason Case 1 worked is social, not technical. A dependency suggested by a colleague gets scrutiny. A dependency suggested by the IDE's assistant gets accepted, because developers have been told for two years that this is what the tool is for. The attacker did not need to defeat code review; they needed to be upstream of the moment when a human decided to trust. This is the same shape as Case 7 and Case 5: in each, the model is a component that converts untrusted input into trusted-looking output. Our assessment: the durable mitigation is not better model judgment but a verification step that does not involve the model — checksums and allowlists enforced at the hook level, as Mandiant recommends, because those hold regardless of how convincing the suggestion was.
Non-determinism is a availability risk, not just a security one. Case 6 deserves more attention than it will get, because nobody attacked anything. An agent with write access to a billing database met malformed data and took the company's transaction processing down while running up a five-figure bill in under an hour. Most organisations have change-control gates for humans touching production databases and none at all for agents doing it. Financial circuit breakers, bounded recursion and rate limits enforced at the service-identity level are unglamorous controls, and they are the ones that would have contained this.
What is genuinely novel, and what is not. Case 4 — malware with embedded lightweight models rewriting its own command strings at runtime to evade detection — is the most forward-looking item in the report, and readers should note it is presented as a mechanism rather than tied to a named victim. Meanwhile the report's own statistic, that vulnerability exploitation has been the top initial access vector for six straight years, is the counterweight to the whole genre. Our assessment: the AI angle in these cases changed the labour cost and tempo of intrusions and created one new trust channel — the assistant's recommendation. It did not replace the fundamentals, and an organisation that cannot patch its externally reachable applications will not be saved by a semantic firewall.
One cross-reference worth flagging: Strix, listed here among the offensive AI tools Mandiant observes, also appears in Gambit Security's reconstruction of an autonomous card-skimming campaign published on 22 September. Two independent teams finding the same open-source harness in unrelated operations is a reasonable indicator that a shared criminal toolchain is consolidating.
What remains unclear
- How the coding-assistant session was hijacked. The report does not say. This is the pivotal step in its headline case, and without it defenders cannot tell whether the entry point was a stolen session token, a malicious extension, a compromised MCP server, or local access to an already-owned workstation.
- When the engagements happened. Most case studies are undated, which makes it impossible to correlate them with public supply-chain events or to judge how current the tradecraft is.
- Victim identities and scale. Expected in IR reporting, but it means the ~100 repositories figure cannot be independently checked.
- The AI-developed 2FA-bypass zero-day. Described as the first publicly confirmed case, dated May 2026, but not linked to a CVE or a disclosure in the material reviewed here. "AI-developed" is doing a lot of work in that sentence and the degree of human involvement is not specified.
- Case 4's real-world footprint. Whether the just-in-time polymorphic malware was recovered from an engagement or assembled as a capability assessment is not clear from the report's framing.
- Whether the Shai-Hulud variant here is the same lineage as the npm worm activity reported earlier in 2026. The report does not make that link explicit.
Lessons and what to do
For security teams. Inventory what your developers' AI tooling can reach — registries, repositories, cloud APIs, internal services — and treat that list as a privileged access review, because that is what it is. Put dependency installation behind a verification hook that checks against an internal registry and an allowlist, so that an assistant's suggestion cannot become an install without passing a check the assistant does not control. Kill long-lived CI/CD credentials in favour of short-lived federated identities; Case 2 exists because one survived. Then add the telemetry: first-time access to a sensitive repository by a service identity is a high-fidelity signal and most SIEMs are not looking for it.
For AI builders. Two controls from this report belong in products rather than in customer runbooks. First, extensions and plugins should never see raw secrets — a broker that issues scoped, short-lived tokens should be the default architecture, not a hardening option. Second, hooks, plugins and MCP servers need signature verification, because Case 3 is simply what happens when an extensibility framework has no integrity layer. Anyone shipping an agent framework should also assume the Case 6 failure mode is theirs to prevent: bounded recursion and cost caps are safety features, not billing features.
For leadership. The uncomfortable item on this list is Case 6, because it required no adversary and would not have been prevented by any security budget. Ask a direct question: which agents in this company can write to a production database, and what stops one from doing so 15,000 times in an hour? Then ask who approved that access and under which change-control process. Mandiant's governance split — policy for AI at the organisational level, technical guardrails of AI at the runtime level — is a useful structure, but the first deliverable is much simpler than a framework. It is a list of agents, their credentials, and their blast radius.
Sources
Primary
- Mandiant / Google Cloud — AI Risk and Resilience Report 2026, September 2026
- Gambit Security — Autonomous AI agents against online retailers, 22 September 2026, cited for the Strix cross-reference
Coverage:
- The Hacker News — Attacker Hijacks AI Coding Assistant Session, Spreads Shai-Hulud Across About 100 Repositories, 16 September 2026
- iTWire — Attackers turn AI coding tools and agent skills into supply-chain entry points
- SecurityBrief — Mandiant warns of AI agents fuelling new attack risks