⌁AI·CYBER·BRIEF▌
Vulnerabilities8 min read

Mooncake, the KV cache layer under vLLM and SGLang, has two unauthenticated flaws — one unfixed

Two CVEs published October 1, 2026 in Mooncake's transfer engine. One lets anyone on the network read and write process memory; the other has no released fix.

Two vulnerabilities published on October 1, 2026 affect Mooncake, the KV cache transfer layer that vLLM, SGLang, TensorRT-LLM and LMCache use to move cached model state between servers. CVE-2026-103764 let anyone who could reach a serving node's data port read and write that process's memory; it is fixed in Mooncake 0.3.13. CVE-2026-103765 leaves the metadata server open to unauthenticated writes, and as of October 3, 2026 no released version fixes it. Upgrade, then make sure neither port is reachable from anywhere you would not trust with the contents of your prompts.

What happened, in plain English

When a large language model answers a question it builds a working scratchpad called a KV cache (key-value cache) — a compressed record of everything it has read so far. Recomputing it for every request is expensive, so modern serving stacks split the work across machines: one pool of GPUs reads the prompt ("prefill"), another generates the answer ("decode"), and the scratchpad is shipped between them over the network. Mooncake, built by Moonshot AI as the serving platform for its Kimi assistant and open-sourced under the kvcache-ai organisation, is one of the most widely used pieces of plumbing for that handoff.

Mooncake's transfer engine has two network-facing parts: a data port that accepts the actual memory transfers, and a metadata server, a small directory service where each node registers a "segment descriptor" — essentially a note saying here is my address, here is the port to reach me on, here is the region of memory you should write into.

Both were built on the assumption that everything talking to them is already trusted. In the first flaw, the data port accepted a short request header containing a memory address and a length, and used that address directly, without checking it against the memory the node had actually registered for transfers — a warehouse that accepts delivery slips with any shelf number written on them, including shelves in the manager's office. In the second, the metadata directory accepted unauthenticated reads, writes and deletions, so the note saying where a cache transfer should go could simply be rewritten.

Both were reported by Mingkai Yu and Jiajia Liu, who filed public issues #4441 and #4444 on October 1, 2026 and demonstrated both against Mooncake 0.3.11.post1.

Are you affected? What to do now

Most organisations are not running Mooncake. It appears in multi-node, high-throughput LLM serving, not in a single-box Ollama or a managed API. But you may be running it without having installed it by name, because vLLM and SGLang pull it in as a KV transfer backend.

Checklist

  • Check whether it is present at all. Run pip show mooncake-transfer-engine in the environment your inference server runs in. On vLLM, also check your launch arguments for --kv-transfer-config containing MooncakeConnector, ECMooncakeConnector or MooncakeStoreConnector; on SGLang, for a Mooncake KV transfer backend. If none appear, you have nothing to do.
  • Check your version. CVE-2026-103764 affects v0.2.0 through 0.3.12.post1. Upgrade to 0.3.13 or later — 0.3.13 was released on August 26, 2026 and the current release is 0.3.13.post1 (August 31, 2026).
  • Treat CVE-2026-103765 as unfixed. It is recorded as affecting Mooncake through 0.3.13.post1 — the newest release. The finder states it still reproduces on 0.3.13 and on the current main branch. Upgrading does not close it.
  • Isolate the metadata server. Find how your config sets metadata_server — in integrations it looks like http://<host>:8080/metadata — and confirm that port is bound to a private interface and firewalled to the serving nodes only. The HTTP metadata backend is the one enabled by default in Mooncake's build (USE_HTTP=ON; etcd and Redis are off by default). If you run the master service, check whether you start it with --enable_http_metadata_server=1.
  • Isolate the data ports too. The transfer engine's TCP data port is chosen from a range (roughly 15000–17000 in the reported configuration) and the master service listens on 50051. None of these should be reachable from a general-purpose VLAN, a VPN pool, a CI runner or anything internet-facing. If your build supports it, the etcd or Redis metadata backend moves the directory off the unauthenticated HTTP handler.
  • Do not expect an automated alert. As of October 3, 2026 neither CVE has a GitHub Security Advisory, so Dependabot and tools that depend on GHSA data will not flag them. You have to check by hand.
  • No indicators of compromise have been published by the project or by either CVE assigner, and there is no public evidence of exploitation in either direction. If you want to look anyway, unexpected PUT or DELETE requests to /metadata, and segment descriptors pointing at hosts outside your cluster, are the things that would matter.

The expert view

The root cause in both cases is the same, and it is the dominant pattern in AI infrastructure security right now: a component designed for a trusted datacentre fabric, shipped as a pip install, deployed by people who reasonably assume a dependency of their inference server is internal plumbing rather than a listening service.

CVE-2026-103764 is CWE-822, untrusted pointer dereference. The transfer engine's value proposition is zero-copy movement of registered memory regions, so the fast path deliberately avoids indirection — and the bounds check that would have made the registration meaningful was absent from ServerSession::readHeader. VulnCheck, the CNA, scores it CVSS 4.0 9.3 (critical) with an AV:N network vector; other trackers carry CVSS 3.1 9.8. The finder's own suggested score in issue #4441 was CVSS 3.1 8.8 with AV:A — adjacent network — on the view that this port belongs on a cluster fabric. Which is right depends entirely on your network: on a flat network it is a 9.8, on a properly segmented fabric it is a serious hardening gap rather than an emergency. The fix is attributed to a substantial restructuring of the TCP transport in 0.3.13 that constrains access to registered buffers; both the CNA and the finder place the remediation in that release.

CVE-2026-103765 is CWE-306, missing authentication for a critical function, scored CVSS 4.0 8.8 (high) by VulnCheck and CVSS 3.1 9.4 (critical) elsewhere. Conceptually it is the more interesting of the two, because it requires no memory corruption at all — it attacks the control plane. If an attacker can rewrite the note saying where a cache transfer should go, the transfer goes there. The detail that should worry operators is the finder's observation that in his test the sending node reported the transfer as successful while the data arrived at his listener. A redirection that returns success is a silent one, and KV cache contents are not an abstract asset: they reconstruct the prompts, retrieved documents and system instructions the model has been processing.

There is also a reporting-quality problem worth flagging, because it will mislead automated tooling. The machine-readable OSV record for CVE-2026-103765 names commit 7197358 as the fix. That commit is titled "Bump version to 0.3.13.post1 in pyproject.toml" and changes a single version string in a single file. Anything reading that field literally concludes the flaw is fixed in 0.3.13.post1 — precisely the release the finder says is still vulnerable, and the version the CVE description itself calls affected. Derived "fixed in" fields in vulnerability feeds are not vendor statements, and should be checked against a project's actual release history before they are trusted.

The wider trend is that the KV cache has become a distinct, poorly governed asset class. It sits outside almost every data classification scheme — not a database, not a log, not a model artefact — while holding a derived copy of the most sensitive text passing through the system. Mooncake's advisory history already runs to eight earlier entries in the GitHub database, and the adjacent integrations have their own: CVE-2026-94627 in the vLLM/Mooncake connector through 0.29.0 and CVE-2026-93688 in SGLang through 0.5.19, both high severity and both from September 2026.

Still unknown: the project has published no security advisory of its own, and both issues were open with no maintainer response as of October 3, 2026, so there is no official statement on the metadata-server flaw or a timeline for a fixed release. It is also unclear whether the 0.3.13 address validation covers transport paths beyond TCP, since the published proof of concept was scoped to TCP. And with no project telemetry and no IOCs, nobody can say whether either flaw has been used against a real deployment.

Official sources

Get the daily brief

AI + security signal by email: headlines, a two-line summary, a link. No noise, no spam.

How often