⌁AI·CYBER·BRIEF▌
Vulnerabilities7 min read

MLflow: two flaws let a crafted model run code even with pickle loading off

CERT/CC published two CVEs in MLflow's dspy and statsmodels model loaders. The statsmodels fix shipped in 3.15.0; the dspy gap is still open in the current release.

On September 23, 2026 the CERT Coordination Center published two vulnerabilities in MLflow, the machine-learning tracking and model-registry platform used by most enterprise data science teams. In both cases a crafted model artifact can execute code on the machine that loads it — even when MLflow's own "do not unpickle untrusted models" safety switch is turned off. One is already fixed (upgrade to MLflow 3.15.0 or later); the other still has no patch in the current release, and the workaround is to stop loading models through the affected component.

What happened, in plain English

MLflow is open-source software that teams use to keep track of machine-learning experiments and to store trained models so that other systems can load and run them. Think of it as a shared warehouse for models: someone puts a model on a shelf, someone else takes it down and uses it.

Models are saved in formats called flavors — one per framework, so mlflow.sklearn, mlflow.pytorch, mlflow.dspy, mlflow.statsmodels and so on. Many of those flavors store the model using pickle, Python's built-in way of writing an object to a file. Pickle has a long-known property that matters here: loading a pickle file does not merely read data, it can run instructions embedded in the file. A malicious model file is therefore closer to an executable than to a document.

Because of that, MLflow added a safety switch: an environment variable called MLFLOW_ALLOW_PICKLE_DESERIALIZATION. Set it to False and MLflow is supposed to refuse to unpickle anything, across all flavors. That is the control many organizations rely on when they load models they did not build themselves.

Researcher Prasanna Dabi found that two flavors do not honour the switch. The statsmodels loader never checks it at all. The dspy loader checks it only when the model file's name ends in .pkl — a file that is a pickle but is named something else takes a different code path and is loaded anyway. In both cases, an attacker who can place a file in a model store that someone will later load gets code execution on the loading machine, with the privileges of whatever process loaded it — typically a training node, a CI runner or an inference server.

Are you affected? What to do now

You are affected if your organisation runs MLflow and loads models it did not produce end-to-end itself — from a shared registry, an artifact bucket several teams can write to, a partner, or a public source.

  • Check your version. Run pip show mlflow (or mlflow --version) on every machine that loads models — tracking server, training nodes, CI runners, inference containers, and developer laptops. Do not assume one number covers all of them.
  • Upgrade to MLflow 3.15.0 or later. This closes CVE-2026-96804, the statsmodels flavor. The fix shipped on July 31, 2026 in release 3.15.0 ("Add MLFLOW_ALLOW_PICKLE_DESERIALIZATION guard to mlflow.statsmodels flavor", PR #24686), well before the CVE was published. Versions 2.1.0 through 3.14.x are affected. The current release at the time of writing is 3.16.1.
  • Stop loading models through the dspy flavor until MLflow ships a fix. CVE-2026-96775 affects MLflow 2.0 and later, and the conditional check is still present in the 3.16.1 source. CERT/CC's advice is explicit: avoid loading any model via the dspy flavor if you are relying on the pickle switch to protect you. If you do not use DSPy, you have nothing to do here.
  • Do not treat MLFLOW_ALLOW_PICKLE_DESERIALIZATION=False as a boundary. It is a useful default, but these two CVEs — and two further advisories fixed in the mlflow.shap flavor on September 3, 2026 — show it is enforced flavor by flavor, not centrally.
  • Check who can write to your model store. The realistic attack path is not the network, it is the shelf. List the identities with write access to your artifact bucket or registry, and remove any that do not need it. An artifact store that is writable by more people than it is readable by is the problematic shape.
  • Isolate the loaders. Model loading should not run as root, should not hold long-lived cloud credentials, and should not sit on the same host as your tracking server if you can avoid it.
  • If you cannot upgrade immediately, restrict loading to models your own pipelines produced and whose artifacts you can account for, and prefer non-pickle serialisation formats where your framework supports them.

There are no indicators of compromise published for either CVE, and no public proof-of-concept had been detected at the time of writing. Note also that MLflow's default deployment has no authentication in front of the tracking server, so "who can write a model artifact" is often a wider set of people than teams assume.

The expert view

Neither of these is an exotic bug. Both are the same failure mode: a security control implemented at the leaf rather than at the trunk.

MLFLOW_ALLOW_PICKLE_DESERIALIZATION reads like a global kill switch, and that is how operators use it. Architecturally it is not one — it is a convention that each flavor's _load_model() is expected to implement for itself. MLflow ships dozens of flavors. Every one of them is a separate opportunity to forget the check, or to implement it slightly wrong. statsmodels forgot it entirely; dspy tied it to a filename extension, which means the control depends on an attribute the attacker chooses. The September 3 mlflow.shap fix and a further pull request hardening pt2 archive validation make three or four instances of the same omission inside a month. That is a pattern, not a coincidence, and it suggests the remaining flavors deserve an audit rather than a wait for the next CVE.

The attack chain is short and does not involve the network stack at all: write a file to a location a victim's pipeline will read, wait for the pipeline to read it. Everything interesting happens inside a deserialisation routine that was designed to reconstruct arbitrary Python objects. This is why the CVSS 3.1 vector for the statsmodels flaw is AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H — 8.8, High. The UI:R is doing real work there: something has to load the model. In a mature MLOps setup, "something loads the model" is a scheduled job that runs several times a day, so in practice that user interaction is closer to automatic than the vector implies. Read the other way, an organisation that only ever loads models its own pipeline produced, from a store nobody else can write to, is not meaningfully exposed. The score is roughly right; the variance between environments is what it fails to capture.

The broader trend is that the ML supply chain is now a routine target rather than a research curiosity. MLflow specifically has been under active attention: CVE-2026-64849, a webhook server-side request forgery flaw fixed in the same 3.15.0 release, was added to CISA's Known Exploited Vulnerabilities catalog on August 19, 2026, and security firm watchTowr reported on August 20, 2026 that attackers began scanning for exposed MLflow instances within hours of the CVE identifier being assigned, exfiltrating cloud credentials. Anyone still below 3.15.0 is carrying both problems.

What remains unknown: MLflow had not issued a statement to CERT/CC as of the note's last update on September 23, 2026, and there is no published timeline for a dspy fix. There is no evidence either CVE has been exploited. One practical caveat for anyone cross-referencing advisories — the CVE-to-flavor mapping is reported inconsistently across trackers; the CVE records themselves are the authority, and they read CVE-2026-96775 as dspy and CVE-2026-96804 as statsmodels.

Official sources

Get the daily brief

AI + security signal by email: headlines, a two-line summary, a link. No noise, no spam.

How often