Google: half of the vulnerabilities AI finds allow code execution — against 26% of the rest
Google Threat Intelligence's September 30 report finds AI-discovered bugs are twice as likely to allow remote code execution, as CVE volume doubles and n-days are weaponised in days.
Google Threat Intelligence Group published research on September 30, 2026 showing that vulnerabilities found by AI tools are about twice as likely to allow remote code execution as vulnerabilities found any other way — 50% against 26%. The same report records monthly CVE disclosures doubling during 2026 and attackers weaponising newly published flaws within days, in one case four days. Nobody needs to patch anything new today because of this report, but anyone who runs a patching programme should read it as a warning that "patch everything" is no longer a workable plan.
What happened, in plain English
A CVE (Common Vulnerabilities and Exposures) identifier is the public reference number a software flaw gets when it is disclosed. Google Threat Intelligence Group — GTIG, the threat research team that includes Mandiant — tracked every CVE disclosed between January 1, 2025 and August 31, 2026, and separately tracked which ones attackers were actually seen using. Then it asked a question nobody had data for before: do the bugs found by AI tools look different from the bugs found by people?
They do. GTIG found that half of the vulnerabilities it could identify as likely discovered by AI lead to remote code execution — the worst common outcome, where an attacker gets to run their own commands on your machine. For vulnerabilities found by other means, that figure is 26%. The AI-found set contains proportionally fewer of the low-value findings: information disclosure drops from 18% to 8%, data manipulation from 9% to 5%.
The reason is mundane rather than magical. Automated agents are good at following long, tedious chains of logic through C and C++ code and at building test cases that prove a crash is real. That is precisely the work that surfaces memory-corruption bugs — and memory-corruption bugs are the ones that turn into code execution.
Meanwhile the raw volume is climbing fast. Disclosures went from 5,045 in January 2026 to a peak of 10,740 in August. GTIG's counter-point matters just as much: only 0.23% of everything disclosed in 2026 — roughly 1 in 431 — was ever observed being exploited. The problem is not that there are more dangerous bugs. The problem is that the dangerous ones are now buried in twice as much noise, and the clock on them is shorter.
Are you affected? What to do now
If you do not own patching or vulnerability triage, there is nothing for you to do today. This is a trends report, not an advisory: there is no new CVE here, no indicator of compromise, no emergency update.
If you do own that work, the report is an argument for changing how you prioritise. A checklist:
- Check one specific thing first. CVE-2026-1731, the BeyondTrust flaw GTIG uses as its case study, was disclosed on February 6, 2026 and fixed in Remote Support 25.3.2 and Privileged Remote Access 25.1.1 (Censys advisory, citing BeyondTrust advisory BT26-02, CVSSv4 9.9). If you run either product and are below those versions, that is an eight-month-old unpatched pre-authentication code execution bug on a remote-access appliance. Fix that before reading further.
- Rank your edge appliances above your servers. Edge and security appliances accounted for 14% of all exploited vulnerabilities in the period; enterprise directory and collaboration systems for 11%. VPN concentrators, firewalls, remote-access gateways, mail gateways — the boxes that are reachable from the internet and that your EDR agent does not run on.
- Close any management interface that faces the internet. More than 65% of the exploited edge flaws targeted unauthenticated public management interfaces. Put the admin panel behind the VPN or an allowlist. This single change removes a whole class of exposure and costs nothing.
- Shorten your n-day window, not your zero-day window. GTIG's finding is that the growth in exploitation comes mostly from fast weaponisation of already disclosed flaws, not from new zero-days. If your service level for a critical internet-facing patch is 30 days, it is being outrun; four days is the observed figure in the BeyondTrust case, with five more attacker clusters joining inside seven.
- Stop treating disclosure volume as a workload. Around 5,000 of the 2026 CVEs came from Linux kernel descriptions alone, with zero observed in-the-wild zero-day exploitation among them. Filter by exploitation evidence (CISA's Known Exploited Vulnerabilities catalogue, EPSS, your threat-intel feed) before filtering by severity score.
- If you run AI infrastructure, inventory it like infrastructure. GTIG counted 2,076 AI-related CVEs across the 20-month window, over 1,500 of them in 2026, and orchestration middleware alone — LangChain, Langflow, Flowise, Dify, MCP servers and similar — accounts for half. It names active in-the-wild exploitation of CVE-2026-42271 in LiteLLM's MCP preview endpoint and CVE-2026-5027 in Langflow, and notes CVE-2025-3248 (Langflow) is in CISA's KEV catalogue. Most organisations have no asset record for these services at all.
- Ask suppliers one question. Whether they run AI-assisted code review before release. GTIG's position is that this should become standard practice; it is a reasonable thing to put in a security questionnaire.
The expert view
The genuinely new contribution here is the risk-profile comparison, and it is more careful than the headline suggests. On GTIG's own risk ratings — not CVSS — AI-discovered vulnerabilities do not come out dramatically more severe: High sits at 4% against 3%. What shifts is the middle of the distribution, from Low (69% down to 39%) into Medium (28% up to 58%). The striking number is the consequence class, not the severity band. Agents are not finding scarier bugs so much as they are finding fewer worthless ones, and the ones they find happen to sit in the category attackers care about.
That is a selection effect worth naming. Agentic discovery is being pointed at memory-unsafe code and at parsing and serialisation paths, because those have what Mandiant's companion blueprint calls a binary oracle: the program either crashes or it does not, so success is machine-checkable. Business-logic and authorisation flaws have no such oracle, which is why agents are not finding them and why the comparison is not quite apples to apples. Expect the RCE skew to persist as long as that targeting holds.
The second finding is the one with operational teeth, and it points away from AI as the villain. If zero-day exploitation is only up from 8 to 11 per month on average while total exploitation jumped to 141 distinct flaws in eight months — past the 127 for all of 2025 — then the acceleration is in n-day turnaround. GTIG's reading, offered as a possibility rather than a conclusion, is that language models are being used to diff patches and read advisories and proof-of-concept code into working exploits faster, which is a far lower bar than original bug discovery. Since May 2026 exploitation growth (+127%) has tracked disclosure growth (+128%) almost exactly. That is the number to watch: when the lines diverge upward, something has changed.
Two honest gaps. First, GTIG says plainly that AI attribution is undercounted — there is no standard CVE metadata field for it, cloud and SaaS vendors fix AI-surfaced bugs in production without ever requesting a CVE, and coordinated-disclosure embargoes hide the rest. The real AI-discovered population is larger than the measured one, and nobody knows by how much, which makes the 50/26 split directionally interesting and not a stable statistic. Second, the defensive answer on offer is itself unproven at scale: GTIG points to CodeMender and to the blueprint Mandiant published on July 16, 2026, which is honest about the failure modes — hallucinated patches that break architectural logic, silent false negatives on anything outside training data, alarm fatigue, and runaway compute cost. Its design rule is the sound part: let the agent produce a reproducible test harness, execute it in a sandbox with hard timeouts, discard anything that fails, and send only the survivors to a human. Automate the proving; keep the judging.
The uncomfortable symmetry is that the same capability sits on both sides. The BeyondTrust bug was found by a defensive research agent — Hacktron, whose agent we covered on September 21 — and then exploited in four days. Finding bugs faster only helps if the patching gets faster too, and that half of the equation is still a human organisational problem.
Official sources
- Vulnerability Discovery and Exploitation Trends in the AI Era — Google Threat Intelligence Group, September 30, 2026 (Robin Grunewald, Supriya Mazumdar, Kelli Vanderlee)
- A Blueprint for AI-Assisted Vulnerability Management — Mandiant / Google Cloud, July 16, 2026 (Jules Czarniak)
- CVE-2026-1731 advisory — Censys, February 2026, citing BeyondTrust advisory BT26-02
- BeyondTrust security advisories — vendor advisory index
- CISA Known Exploited Vulnerabilities Catalog — the exploitation-evidence filter referenced above
- Coverage: The vulnerabilities AI finds are the ones attackers want — Help Net Security, October 1, 2026