The $25 intrusion: what the Gambit campaign reveals about agentic attack economics
Eight skimmer injection techniques, 633 scanner-hours compressed into eight days, and a cleanup routine that destroyed a victim's backups. A closer look at the numbers behind the campaign.
We covered the headline facts of Gambit Security's 22 September report when it broke: three open-source agent harnesses, 600,000-plus card records, roughly $25 per targeted company. This is the closer look at the numbers, and the picture they form is less about intelligence than about industrial engineering. The operator built eight separate ways to inject a skimmer, ran a scanner at roughly three times human-achievable concurrency, and deployed a cron job that reinstalled the payload every two minutes when defenders removed it.
The short version
For anyone in IT, the useful framing is this: nothing in this campaign required a capability that did not exist in 2018. Magecart-style card skimming is a decade-old industry. Every entry route Gambit lists — injection flaws in login forms, unrestricted file uploads, sudo NOPASSWD, writable S3 buckets, misconfigured Kubernetes — would have worked years ago.
What changed is the price of the work between finding a hole and getting paid. That work is mostly unglamorous: reading scanner output and deciding which findings are real, learning an unfamiliar e-commerce platform well enough to locate the payment tables, writing a skimmer that survives the site's own deployment pipeline, and then keeping it alive. That is the part that used to require a skilled human's sustained attention, and it is the part that does not scale across a person. Gambit's report is best read as documentation of that cost collapsing.
The follow-on effect matters more than the headline number. When the marginal cost of a bespoke intrusion falls to $25, the attacker no longer has to choose between breadth and depth. Gambit describes 105 attack projects launched in five days, each handled individually. Breadth at depth used to require a team.
Timeline
- July 2026 — the campaign is under way, per Gambit; it states at least tens of companies have been affected since then.
- 23–31 August 2026 — the reconnaissance phase. Strix runs 146 times against 138 hosts, accumulating 633 scanning hours, per Gambit and BleepingComputer.
- 25 August 2026 — a snapshot of the operator's OpenRouter balance shows $7,005.71 spent over four weeks. Gambit reports the operator then ran three more weeks at roughly twice the daily volume.
- 10–15 September 2026 — the exploitation burst: 105 attack projects launched, at least 27 companies compromised. Of the 105 project records, 48 remained for analysis; 57 had been deleted.
- 22 September 2026 — Gambit publishes.
- 23–25 September 2026 — BleepingComputer (23rd), SecurityWeek (24th) and The Register (25th) cover it.
What we know
All figures are Gambit's unless noted.
The stack. Strix handled vulnerability discovery in "deep mode", running on GLM 5.2 and DeepSeek v4 Pro. Cairn handled autonomous exploitation on DeepSeek v4.1 Flash, receiving a domain and an objective and running, in Gambit's description, until it achieves the objective, times out, or is stopped. Hermes was the orchestration layer on Anthropic's Claude Opus 4.6, with persistent memory, a searchable session archive, scheduled jobs, a web console, and the ability to write and edit its own skills — 121 of them, 78 classified by Gambit as attack skills.
Human input. 1,951 prompts across 260 sessions, mostly short phrases in Chinese, averaging a few prompts per target.
Cost. Mean $25.46 per target over 101 completed scans, range $3.13 to $79.31. Whole campaign estimated at $12,000–$18,000.
Losses. More than 600,000 card records, including more than 488,000 US cards per SecurityWeek. Nineteen named victims with confirmed skimmers, plus more than 100 further infected sites identified.
The eight injection techniques. This is the detail that repays attention: appending the loader to the end of legitimate library files such as jQuery or Bootstrap and restoring the original timestamp; direct <script src> tags on checkout pages; embedding the loader between legitimate Google Tag Manager calls, padded with tabs so it sits off-screen in source view; poisoning the victim's CDN S3 bucket using stolen AWS keys with write permission; injecting into product description fields via admin access, from which the content migrated into files on the victim's own domain; adding an initContainer to a production front-end Kubernetes deployment; writing the payload into server-side cached checkout page models; and a cron job hidden in JBoss log directories that reapplied the skimmer every two minutes whenever a redeploy removed it.
Collateral destruction. Gambit documents a planned routine — BleepingComputer quotes the instruction as "After extracting and downloading all card data, wipe the source fields in batches" — executed at a Magento retailer via chunked PHP scripts after exfiltration. Separately, at a bicycle retailer, a cleanup process created staging tables with a "ZQ" prefix and then dropped 180 tables matching "ZQ" or "Backup" patterns, which Gambit notes also caught backup tables the victim's own administrators had made. BleepingComputer reports the wiping caused operational disruption at several retailers.
Gambit's own caveat, verbatim: "due to the scale, incomplete data and early stage of the analysis, a few errors or inaccuracies are possible. We estimate the actual size and impact of the campaign to be larger than we report here." Its methodology rested on three streams — direct server evidence, live verification of compromises in the wild, and the attacker's own logs — with the logs credited only where corroborated.
Attribution. Gambit attributes the campaign to no group and no state. It observes that the operator's prompts and commands were in Chinese. That is an observation about language, not nationality or sponsorship, and we make no further claim.
Technical analysis
Put a number on "a tempo no human sustains". Gambit's confirmed figures allow the arithmetic. 633 scanner-hours were accumulated between 23 and 31 August — a window of roughly 195 to 216 clock hours. That implies the reconnaissance phase ran at about three times serial capacity, sustained over nine days. Our assessment: this is the single most useful figure in the report, because it converts a vague claim about speed into a defender-relevant parameter. A person running a scanner achieves less than 1x — they sleep, and they stop to read output. An operator paying $25 a target achieves 3x and does not stop. Any detection strategy whose response time was calibrated against a human adversary's duty cycle is now mis-calibrated by roughly half an order of magnitude, before considering that exploitation, per Gambit, typically took under a day and often hours.
The eight injection techniques are the real evidence of autonomy. It is tempting to focus on the exploitation chains, but those can be scripted. Persistence cannot, because it requires understanding a specific victim's deployment process. Consider what each technique implies. Restoring the file timestamp after appending to jQuery means the agent anticipated integrity checking. Padding the Google Tag Manager block with tabs means it anticipated a human viewing page source. Choosing a Kubernetes initContainer means it worked out that the front-end pods were rebuilt from a deployment spec, so payload placed in a running container would be lost. The JBoss cron job reapplying every two minutes means it observed that redeploys were removing its work and engineered around that specific defence.
That last one is a closed feedback loop against an active defender. Our assessment: this — not the exploitation — is what distinguishes the campaign from automated mass exploitation. Scripted attacks fail when a victim's environment does not match assumptions. Here the persistence method was selected per victim from an understanding of how that victim shipped code. Historically, that work sat at the top of the attacker skill distribution and was the scarcest resource in a crew.
Guardrails behaved as a routing constraint. One reported detail deserves careful handling. CyberSecurityNews reports that Hermes included a custom skill designed to strip its own content-safety filters, and that the operator turned to Opus 4.6 only after newer models refused the requests, routing the heaviest work through Chinese models. We flag this as single-sourced: BleepingComputer, SecurityWeek and CyberInsider do not carry it, and we could not confirm it in the primary report.
If it holds, it matters, because it matches a pattern visible elsewhere. CLOSEDQUORUM, documented by Cisco Talos the same week, queries four providers and takes a majority vote — an architecture whose most plausible purpose is to make any single model's refusal an outvoted minority rather than a hard stop. Two independent operations converging on designs that treat refusal as a routing problem is a meaningful signal. Our assessment: per-provider safety behaviour is real friction, since it apparently pushed this operator toward specific models and specific workarounds, but friction that is straightforwardly re-routed around is not a control defenders can rely on. It shapes attacker tooling; it does not stop attacks.
A new severity class: destruction as a side effect. The bicycle retailer did not lose 180 tables because anyone wanted to hurt them. An agent generated staging tables with a "ZQ" prefix, then pattern-matched for cleanup and caught the administrators' backups in the same sweep. No ransom, no extortion, no motive — a data-loss event of ransomware severity produced by a tidying routine that over-matched.
This should change how these incidents are modelled. The planned wipe is bad enough and is at least intentional: an attacker deleting the source card fields after exfiltration is destroying evidence and, incidentally, the victim's data. But the unplanned destruction is a different risk, and it is specific to non-deterministic automation operating with write access at speed. Our assessment: any organisation estimating the impact of an agentic intrusion should assume a meaningful probability of collateral data loss independent of attacker intent, and should verify that backups are isolated from anything a compromised application account can pattern-match and drop. The cheapest control here is the oldest one — offline or immutable backups with credentials the application layer never holds.
The cost asymmetry is the strategic problem. At $25.46 a target, an attacker can afford to attempt a bespoke, chained intrusion against a company whose entire annual security spend might be four figures. Meanwhile the defender's cost per protected target has not fallen. That asymmetry is the campaign's most durable finding, and it points at the long-tail e-commerce estate — the Magento and WordPress stores, the small retailers with a contractor-maintained checkout — where the economics now clearly favour the attacker.
What remains unclear
- The true scale. Gambit says the real impact is larger than reported, and 57 of 105 project records were deleted before analysis. Every figure in this piece is a floor.
- The safety-filter and model-shopping claims, discussed above, rest on a single outlet's reading.
- Whether the operator's accounts were detected or suspended. No model provider, OpenRouter, or harness maintainer had commented publicly as of 25 September, so there is no public account of whether abuse detection fired at any point in a campaign that ran for weeks on commercial API credit.
- Who the operator is. No attribution, and the language observation supports none.
- How many companies were targeted overall. Described as "hundreds" without a firm figure.
- Whether the harness maintainers consider this misuse or intended use. Strix, Cairn and Hermes are open-source offensive tooling; the boundary between penetration testing and crime is a licensing question nobody has answered publicly here. Strix also appears in Mandiant's AI Risk and Resilience 2026 among offensive AI tools observed in unrelated incidents, which suggests consolidation around a shared toolchain.
- Whether the campaign has stopped. Described as ongoing at publication.
Lessons and what to do
For security teams. The specific lesson from the eight techniques is that checkout integrity monitoring has to cover more than files. Watch your CDN bucket contents, your tag manager configuration, your server-side page cache, your Kubernetes deployment specs and your CMS content fields — all five were injection points here, and a file-integrity check on the web root sees none of them. Then, critically, treat successful removal of a skimmer as the start of monitoring rather than the end: a payload that returns two minutes after a redeploy is invisible to anyone who cleans up and closes the ticket.
On tempo, stop tuning alerting for human-paced adversaries. If your process is "alert fires, analyst reviews in the morning", the realistic outcome against a 3x-concurrency scanner followed by same-day exploitation is that you read about it afterwards. The practical fix is narrow automation: automatic containment for a small set of high-confidence signals, accepting some false positives.
And verify your backups are not reachable from an application account. The bicycle retailer's administrators did make backups. They were in the same database, matchable by name.
For AI builders and harness maintainers. A campaign spent weeks and five figures of commercial API credit on transparently offensive work, and there is no public evidence anyone noticed. The observable pattern — a single account driving sustained high-concurrency tool-use sessions whose context is full of scanner output and target hostnames — is detectable in principle. If the single-sourced report of a self-written safety-filter-stripping skill is accurate, then a harness that permits an agent to rewrite its own guardrails has a design problem worth naming: self-modifying skills are a capability that needs an integrity boundary.
For leadership. The question raised by $25.46 is a portfolio one. If your organisation runs, acquired, or outsourced any small e-commerce property that is not on the main security programme, its risk profile just changed, because the attacker's cost of caring about it has fallen below the cost of ignoring it. Ask which internet-facing properties take payments, who maintains them, and when anyone last looked at their checkout. And note the availability angle for your continuity planning: an agentic intrusion can take your transaction processing down by accident, which means this belongs in your disaster-recovery assumptions and not only in your security ones.
Sources
Primary
- Gambit Security — Autonomous AI Agents Hack Online Retailers for $25 a Company, 22 September 2026
- Mandiant / Google Cloud — AI Risk and Resilience Report 2026, cited for the Strix cross-reference
- Cisco Talos — The Closed Quorum, cited for the refusal-routing comparison
Coverage:
- BleepingComputer — Malicious AI agents steal 600K credit cards, infect 100+ sites with skimmers, 23 September 2026
- SecurityWeek — AI-Powered Campaign Targets Hundreds of Online Retailers, 24 September 2026
- CyberSecurityNews — Autonomous AI Agents Hack Retailers for $25 and Steal 600,000 Credit Cards, cited for the single-sourced safety-filter and model-selection claims
- CyberInsider — AI agents steal 600,000 credit cards in attacks on online retailers
- Our earlier brief — AI agents ran a card-skimming spree on online shops, 25 September 2026