Skip to content
The AI-First Web

Exposed LMCache servers can be taken over with one message; no fix yet

JFrog's CVE-2026-105192 shows how AI infrastructure treats the internal network as a trusted place

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Exposed LMCache servers can be taken over with one message; no fix yet
In brief
  • JFrog found that LMCache's multiprocess mode lets an unauthenticated sender run code when bound to a routable address (CVE-2026-105192). No fixed version exists.
  • This flaw shows how a service built for trusted neighbors can become reachable by anyone once one setting or a sample configuration opens its port.
  • Check whether any LMCache server runs in multiprocess mode on a routable address. Until a fix ships, keep it on localhost or a trusted cluster network.

Much AI infrastructure is built on an unspoken promise: only our own processes will ever talk to this service. The LMCache flaw shows what happens when that promise meets a real network. A service designed for friendly neighbors will answer anyone who knocks.

What JFrog found

LMCache is a shared cache for AI model servers such as vLLM, published as open source. JFrog Security Research disclosed a flaw in its multiprocess mode, also called distributed mode. The flaw is tracked as CVE-2026-105192 and scored 9.8 out of 10.

In this mode, LMCache opens a network socket so worker processes can register and share cached model data. JFrog found no authentication on that socket. There is no password, no encryption handshake and no check on messages. The intended callers are sibling LMCache processes, but the socket does not verify who is calling.

9.8 / 10
Severity score for CVE-2026-105192
Source: JFrog Security Research (October 7, 2026)

How one message becomes code execution

Messages on the socket use a compact format called msgpack. One message type carries a special wrapper, marked as extension code 1. When the server decodes it, the wrapper's data is passed to pickle.loads.

Pickle is a Python format that can carry executable instructions. Unpacking untrusted pickle data hands control to whoever wrote it.

The order of events makes this worse. The server unpacks the data while it is still reading the message's arguments. That happens before the handler runs and before the message is checked. JFrog's proof of concept needs a single message to the default port, 5555.

The server then logs a type error, because it expected a different object. By then the attacker's command has already run. A log that shows only an error would not reveal the attack.

The code runs with the privileges of whoever owns the LMCache process. On the project's ready-made container images, JFrog reports, that owner is root, the most powerful account on the system.

One setting decides your exposure

The limits matter here. By default, the server accepts connections only from its own machine, so other hosts cannot reach it. Teams that run LMCache only inside a single vLLM process never expose this port.

The 9.8 score applies when an operator sets a routable address with the --host option. JFrog says this is how multi-node deployments let peers connect. The Hacker News says the project's own example Kubernetes deployment tells the server to accept connections on every network interface.

So the risk is not in every LMCache install. It sits in deployments bound to a routable address, and multi-node deployments are the typical case for that.

Which versions are affected

The vulnerable decode path first shipped in v0.3.9. It remains in v0.5.5, the latest PyPI release. The 0.5.6 release candidates, up to rc3, carry the same code, and so does the dev branch. JFrog confirmed this on October 7, 2026.

The 0.5.6 release candidates are not a fix. An operator who upgrades to rc3 stays exposed.

v0.3.9 to v0.5.6rc3, plus dev
Affected builds
Source: JFrog Security Research (October 7, 2026); no fixed version published

The trust boundary that was never drawn

The lesson here is that the inside of a cluster is not a security control. A platform engineer who copies the project's example deployment to get multi-node caching working is doing what that example shows. Nothing in that path asks whether the network can be trusted. In our reading, nobody on the team may have decided to accept the risk, yet the risk could still be accepted.

This is also a familiar pattern. The Hacker News ties the flaw to ShadowMQ, a name researchers gave in November 2025 to similar bugs in other AI inference frameworks. In those cases too, data from an unauthenticated socket reached pickle. That is context, not a count. This story does not claim a trend beyond what JFrog reports for LMCache.

The practical cost falls on operators. No patched release exists, and LMCache has not published a security advisory, according to The Hacker News. JFrog's advisory gives no way to check whether a server has already been attacked.

Unconfirmed reports to keep separate

The Hacker News also reports that one GitHub account opened six more LMCache security reports on October 6. They allege cross-tenant data access and network services that run commands without a login.

6
Additional unconfirmed reports from one account
Source: The Hacker News (October 7, 2026); no CVE, no maintainer confirmation, no fix

Treat these as allegations. The same coverage notes that one of them points to a default that changed in the 0.5.6 release candidates: an admin web server now listens only on the local host. That change does not address CVE-2026-105192, since those builds still carry the pickle decode path.

A related vLLM denial-of-service bug, CVE-2026-105756, is rated 6.5. It does not allow code execution and was fixed in vLLM 0.30.0.

What to ask your team this week

Start with inventory. Ask whether any team runs LMCache in multiprocess mode. If so, ask what value --host has and which hosts can open a connection to port 5555. Ask which version is running, and do not accept 0.5.6 release candidates as an answer.

Then ask whether the container runs as root. JFrog's advice is to keep the port on localhost or a trusted cluster network until a fix ships. A firewall narrows who can connect, but any host that can still reach the port can run code.

Also ask who is watching for a fix. With no advisory from the project, someone needs to track new releases and the CVE record directly.

When exposed, a cache server that can be told to run code is not a cache server. It is an open door into the machines that serve your models.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: JFrog Security Research.

Share this insight