⌁AI·CYBER·BRIEF▌
Vulnerabilities8 min read

Ten advisories, no patched version: LightLLM's internal services are open by default

Ten CVEs now list ModelTC's LightLLM inference server as affected, nine of them through 1.2.0 — its newest release. Six are unauthenticated remote code execution. There is no version to upgrade to.

Ten security advisories published between February 17 and September 30, 2026 list ModelTC's LightLLM inference server as affected, nine of them naming every version through 1.2.0 — which is its newest release, shipped on August 10, 2026. Seven are rated critical; six of those are unauthenticated remote code execution in internal services that listen on every network interface by default. There is no patched version to upgrade to, so the only remediation available today is network isolation.

What happened, in plain English

LightLLM is an open-source server for running large language models: you point it at a model and it answers requests over the network. It is not a single program but a group of cooperating processes — a web front end, a router that schedules requests, GPU workers that do the arithmetic, a separate service for images and audio, and several helpers for caching, profiling and moving data between machines.

Those helpers talk to each other over the network too, and that is where the problem sits. Several exchange data using Python's pickle format. Pickle is not really a data format; it is a short set of instructions for rebuilding a Python object, and rebuilding an object can involve running code. Handing an untrusted pickle to a program is close to handing it a script and asking it to run. Several LightLLM helper services accept pickled data from anyone who can reach their port — no password, no check on who is connecting — and by default they listen on all network interfaces rather than only on the local machine.

An analogy: the front door has a receptionist, but the service corridors behind it — goods lift, maintenance hatch, internal phone line — open onto the street, and anyone arriving that way is treated as staff.

Researchers working through the codebase published findings on service after service: the prefill-decode (PD) master node, the config server, the node-registration endpoint, the NCCL data-transfer worker, the router profiler, the multimodal cache, the image service and the reinforcement-learning control routes. ("NCCL" is NVIDIA's library for moving data between GPUs; "RPyC" is a Python library for calling functions on another machine.) Nine of the ten list the same affected range: every version up to and including 1.2.0. The tenth, from February 2026, stops at 1.1.0.

Are you affected? What to do now

You are affected if you run ModelTC LightLLM at version 1.2.0 or earlier, self-hosted in a container or on bare metal, on-premises or in a cloud VM. These advisories do not touch vLLM, SGLang, TGI, Ollama or hosted inference APIs.

One thing to check first: LightLLM is not LiteLLM. They are unrelated projects with near-identical names, and advisory feeds mix them up constantly. These ten CVEs are LightLLM (ModelTC's inference server), not LiteLLM (the AI gateway).

  • Confirm whether you actually run it. pip show lightllm is not reliable here: the lightllm package on PyPI sits at version 0.0.1, last released in August 2024, and is not how this project is distributed. LightLLM is normally installed from the Git repository or a container image, so check your checkout with git describe --tags, or read the image tag of the running container.
  • Do not plan around an upgrade. The newest release is v1.2.0 (August 10, 2026), and nine of the ten advisories list 1.2.0 as affected. GitHub's advisory entries show no patched version. Three of the reports filed in late September were still open with no maintainer response as of October 2, 2026; on an earlier one a contributor offered to take it on, but no pull request is linked.
  • Inventory what is listening. On each inference host, list the open ports (ss -lntp) and compare against what you meant to expose. The affected services bind to 0.0.0.0 by default — every interface on the box, including any public one.
  • Isolate, then isolate again. Only the HTTP inference API should be reachable by clients, ideally through an authenticating reverse proxy. Every other LightLLM port — RPyC channels, config server, PD registration, visual service — belongs on loopback or a private segment with an explicit allow-list. In Kubernetes, a default-deny NetworkPolicy on the inference namespace does most of this work.
  • Switch off what you are not using. The profiler service only runs with --enable_profiling; the NCCL control channel is tied to --pd_trans_mode nccl. Turn both off in production. Note that the reporter of CVE-2026-103270 found the reinforcement-learning control routes mount on the public HTTP app regardless of the --enable_rl flag, so that one cannot be disabled by configuration.
  • If you serve images or audio, constrain outbound traffic from the inference host. CVE-2026-103243 is a server-side request forgery issue: unvalidated image_url and audio_url parameters let an outsider make your server fetch arbitrary URLs, including cloud metadata endpoints.
  • Do not expect automated tooling to tell you. The GitHub advisories are marked unreviewed and LightLLM has no supported package ecosystem, so Dependabot will not raise an alert for any of them.
  • If one of these hosts was reachable from an untrusted network, treat it as a possible compromise rather than a patching exercise: review process and outbound-connection logs, and rotate any API keys, cloud credentials or tokens present in that process's environment. No indicators of compromise have been published by any official source, and there is no public evidence of exploitation in the wild.

If your inference servers already sit on a private subnet with no untrusted tenants and no internet route, you have hardening to do, not an incident.

The expert view

The striking thing is not any single bug but the uniformity. Six of the ten are the same defect in different components: an internal control channel — RPyC in four cases, a raw WebSocket in two — that deserializes attacker-supplied data with pickle enabled, no authentication, bound to all interfaces. CWE-502 across those six; the rest are CWE-306 (missing authentication) on the registration and RL-route issues, CWE-918 (request forgery) on image handling, and one memory-exhaustion bug. VulnCheck, the CNA for the set, scores the critical ones at CVSS 4.0 9.3; several secondary trackers publish CVSS 3.1 figures of 9.8 for the same records, which is the usual disagreement between the two scoring versions rather than a factual dispute.

Conceptually the chain is short: reach the port, send something the service will deserialize, and code runs inside the inference process. That process is a rich target — GPU access, the loaded model weights, whatever provider keys live in its environment, and, in a cloud VM, a route to the instance metadata service. The authentication bypass on /pd_register adds a quieter variant: register a bogus node and the PD master will route real user prompts to it.

The underlying failure is an architectural assumption, not a coding slip: LightLLM is a distributed system whose components were designed as if the network between them were trusted, then packaged as a quickstart people run on a single reachable host. That is a well-worn path — Redis, Memcached, Elasticsearch, Jupyter and Ray each had their own version of it — and the inference layer is where it is repeating, because GPU serving stacks get deployed by teams whose priority is throughput, not perimeter design. The repository has no SECURITY.md.

Where severity is genuinely misread is in the metadata rather than the scores. The machine-readable OSV records for at least four of these CVEs, including the newest, carry a Git range whose fixed value is commit 65c174e. That commit is dated August 4, 2026 — six days before v1.2.0 shipped — is titled "add in-process URL pool caching (#1325)", and changes six files: the CLI argument parser, a start-arguments dataclass, an environment-variable helper, a multimodal utility and two documentation pages. It touches none of the vulnerable services, and it is not the commit the v1.2.0 tag points at (25475e9). A commit that predates the release nine of these advisories call vulnerable cannot be the fix for bugs published weeks after it landed. The OSV entry for CVE-2026-103395 also carries last_affected: 1.2.0 and fixed: 1.2.0 simultaneously. Any tool that reads those fields literally will conclude the issues are resolved in a release that every advisory says is vulnerable. For an operator, the practical lesson is that "fixed in" fields are derived data and need to be read against the vendor's actual release history. It is not an isolated case: the same week, OSV named a one-line version bump as the fix for one of two Mooncake flaws that the newest Mooncake release still carries.

What remains unknown: whether the maintainers intend to address the set, how many instances are exposed to the internet, and whether anyone has attempted exploitation. None of the three has a published answer.

Official sources

Get the daily brief

AI + security signal by email: headlines, a two-line summary, a link. No noise, no spam.

How often