Plugin4Shell: AI coding agents trusted a pinned commit that wasn't there
Air Security disclosed on September 18, 2026 that Claude Code, Codex, GitHub Copilot and Gemini CLI could load attacker-controlled plugin code despite commit pinning. Two are patched, two are not.
Security firm Air Security published research on September 18, 2026 describing a flaw it calls Plugin4Shell, which let four of the most widely used AI coding agents — Anthropic's Claude Code, OpenAI's Codex, GitHub Copilot and Google's Gemini CLI — install plugin code other than the code they had been pinned to. Anthropic and OpenAI shipped fixes months before the public write-up (Claude Code 2.1.179, Codex 0.146.0); according to Air Security and reporting on September 18, GitHub Copilot had no fix and Google declined to patch Gemini CLI, pointing users to a successor product. If your team uses AI coding agents with plugins, check your versions and find out which git servers those plugins come from.
What happened, in plain English
An AI coding agent is a tool that writes and runs code on a developer's machine on request. Most of them support plugins — add-ons, usually pulled straight from a git repository, that extend what the agent can do.
Because plugins run with the developer's privileges, the agents tried to be careful. Instead of always taking the newest version of a plugin, they pinned it: they recorded a commit SHA — the 40-character fingerprint git assigns to one exact snapshot of a repository — and asked git for that snapshot every time. The idea is sound. A SHA is derived from the content itself, so it cannot be quietly pointed at different code.
The problem was in the asking. Air Security found that the agents ran the equivalent of "check out this SHA" and then never verified that the code they ended up with was actually that commit. In git, a 40-character hex string can also be a perfectly legal branch name, and when a name could mean either a branch or a commit, git prefers the branch. So a repository owner who named a branch after the pinned SHA, and made it the repository's default, could serve entirely different code while the pin still looked honoured.
An analogy: ordering a book by its ISBN is precise, because the number describes the book. This was closer to ordering by ISBN and accepting whatever the shop handed over from a shelf someone had relabelled with that number — without ever opening the cover to check.
Air Security describes a second variant affecting Gemini CLI, where the agent fetched the right commit but then checked out the generic reference FETCH_HEAD, which a branch of that name could shadow.
Are you affected? What to do now
You should care if your developers use Claude Code, Codex, GitHub Copilot or Gemini CLI with plugins. If your agents run without plugins, this particular issue does not reach you.
There is an important limiting factor. Exploitation depends on a git host that permits a branch name made of 40 hex characters. GitHub rejects such names, so plugins hosted on github.com are not exploitable this way; per Help Net Security's September 18 report, Bitbucket and self-hosted git servers do allow them.
Checklist, in order:
- Check your agent versions. Claude Code must be 2.1.179 or newer (
claude --version); Codex must be 0.146.0 or newer (codex --version). Both fixes are already released — most teams on auto-update will be past them. - Inventory where your plugins come from. List the plugins and plugin marketplaces your agents are configured to use, and note the git host of each. Anything on Bitbucket, GitLab self-managed, Gitea or another self-hosted server is where attention belongs; github.com-hosted plugins are not exposed to the branch-name trick.
- For GitHub Copilot: as of the September 18 disclosure Microsoft had not shipped a fix and, per Help Net Security, had not issued a statement. Treat Copilot plugins from non-GitHub hosts as untrusted until Microsoft confirms a fix. Verify the current status with Microsoft before relying on this.
- For Gemini CLI: Google confirmed on August 4, 2026 that it would not patch, because the product is deprecated; it directs users to migrate to Antigravity. Plan that migration, or stop using Gemini CLI plugins.
- If you run your own git server: consider rejecting branch and tag names that look like 40-character hex strings, which removes the precondition regardless of which agent your developers use.
- Do not go hunting for indicators. Neither Air Security nor any vendor has published a CVE identifier, a CVSS score or indicators of compromise. There is nothing official to search your logs for, and no public evidence of exploitation in the wild.
For most readers the practical work is small: confirm two version numbers, then answer one question about where your plugins are hosted.
The expert view
The root cause is a category error that recurs across supply-chain tooling: treating a name as if it were a content address. A commit SHA is a content address, and that is exactly why pinning to one is supposed to be safe. But git checkout <string> resolves through git's reference namespace first, so passing a SHA as a string hands the decision back to whatever the remote says that name means. The integrity property lived in the identifier, and the code threw it away at the moment of use.
The fix is correspondingly simple, and visible in OpenAI's changelog for Codex 0.146.0 — "Verify Git plugin SHA checkouts" (pull request #34644): after checking out, confirm that the resulting HEAD commit really is the pinned one. This is verify-after-write, and its absence is the whole bug.
Three things made it worse than a local mistake. First, plugin marketplace review inspects a snapshot at submission time, while delivery happens later and separately — so review and delivery could disagree, and only delivery matters. Second, agents update plugins in the background, so no user action was required; this is what makes "zero-click" a fair description. Third, the trust boundary was invisible to the victim: a pinned reference in a config file is precisely the artefact a reviewer would point at to argue the dependency was safe.
On severity, judge it without a score, because none was published. The blast radius is real — arbitrary code execution in a developer's environment, which typically holds source, cloud credentials and CI access. But the precondition is meaningful: the attacker must control the plugin repository and it must live on a host that tolerates SHA-shaped branch names. Against a marketplace plugin whose maintainer is honest and whose code sits on GitHub, this does not work. Against an internal plugin on a self-hosted server whose access controls are looser than anyone assumes, it plausibly does. Our reading is that the specific bug is bounded, and the pattern it exposes is not.
What remains unknown: whether GitHub Copilot has been fixed since September 18; Air Security's characterisation that "millions of agents" were affected comes with no published methodology and should be read as an order-of-magnitude claim about install bases, not a count of exposed systems; and no vendor has said whether it looked for exploitation before patching. It is also worth noting that Claude Code's public release notes for 2.1.179 (June 16, 2026) list no security entry — the fix shipped silently, which is common practice for coordinated disclosure but means version numbers, not changelogs, are what teams have to go on.
Official sources
- Air Security — Plugin4Shell: Zero Click RCE Vulnerability found in top 4 most popular coding agents (original research and disclosure timeline)
- OpenAI Codex release 0.146.0 — changelog entry "Verify Git plugin SHA checkouts" (#34644)
- Anthropic Claude Code release v2.1.179 — the release containing the fix; no security entry in the public notes
- Coverage: Help Net Security — Zero-click RCE vulnerability hit four major AI coding agents, two remain unpatched
- Coverage: The Hacker News — Plugin4Shell Lets Repository Owners Swap Pinned Plugin Code Across Four AI Coding Agents
- Coverage: The Register — AI coding agents' 0-click RCE flaw could hand attackers keys to the kingdom