Unpatched LMCache Flaw Allows Remote Code Execution
AI-generated image

Unpatched LMCache Flaw Allows Remote Code Execution

A critical unpatched vulnerability in LMCache lets attackers run code on cache servers without authentication. Versions 0.3.9 through 0.5.5 are affected, and no…

A critical unpatched LMCache vulnerability allows an attacker to run code remotely on a cache server without first logging in, according to a disclosure published by JFrog on October 7. LMCache is open source software that speeds up large language model (LLM) servers such as vLLM, tools that generate and process text with artificial intelligence, by storing repeated work so the model does not have to recompute it. The flaw lives in LMCache's multiprocess mode, where the cache runs as a standalone server that LLM workers reach over the ZeroMQ messaging library, a system for passing messages between programs. A single crafted network message to that server can execute commands with the same privileges as the LMCache process.

The ZeroMQ socket that the multiprocess server opens for worker processes to register and share cached data has no authentication, meaning it accepts messages without checking who sent them. One type of message is unpacked with pickle, a Python format that can carry code and run it while the data is being decoded. The server unpacks that message while still reading its arguments, before it checks the message type, so a crafted message can run the sender's code. JFrog's security research team found the flaw, and the discovery is credited to Yuval Moravchick. On the project's official container images, the LMCache process runs as root, the most powerful account on a Linux system, so a successful attack can take full control of the machine, according to JFrog.

Whether a server is reachable from another machine depends on one setting. By default, the multiprocess server listens only on the local machine, so another host cannot reach it. It becomes reachable when an operator starts it with a routable address, as is done in multi-node deployments that share a cache across machines. LMCache's own example Kubernetes deployment, a system for running containerized applications, starts the server that way, listening on every network interface. A copy of LMCache running inside a single vLLM process does not open the port at all.

The vulnerability is tracked as CVE-2026-105192 and carries a severity score of 9.8 out of 10, in the critical range, the rating JFrog gives to a server bound to a routable address. It affects LMCache from version 0.3.9, released in October 2025, through 0.5.5, the latest stable release, and is also present in the 0.5.6 release candidates and the development branch. No fixed version exists. LMCache has not published a security advisory for the flaw, and JFrog's advisory does not offer operators a way to determine whether a server has already been attacked.

Separately, a GitHub user opened six additional LMCache security reports on October 6, the day before CVE-2026-105192 was made public. Those reports allege unauthenticated access to cached data belonging to different tenants, as well as access to several network services that execute commands without a login. They come from one account, rest on proof-of-concept claims, and have no CVE, no confirmation from the maintainers, and no fix. One report points to a default LMCache behavior that has since changed: an admin HTTP server that listened on every network interface in version 0.5.5 listens only on the local host in the 0.5.6 release candidates.

A related flaw in vLLM is already fixed. Before version 0.30.0, released September 22, a single request carrying a malformed cache_salt value could crash the engine on deployments that use the LMCache multiprocess connector. That denial-of-service bug is tracked as CVE-2026-105756, rated 6.5, and does not allow code execution. The core mistake, handing data from an unauthenticated network socket to pickle, is the same one researchers found across other AI inference frameworks in November 2025, in a group of flaws they called ShadowMQ. Whether LMCache's code shares a common source with those projects has not been established.

Until a patched release ships, JFrog advises operators not to assign the multiprocess server a routable address and to keep its port on the local machine or on a trusted cluster network. A firewall that limits who can reach the port lowers the risk but does not remove it, because any host that can still open a connection can run code. For IT teams running AI caching or any public-facing infrastructure, keeping unauthenticated management and cache endpoints off the public internet and limiting access to trusted networks is exactly the kind of control that AEU-I applies as part of security-first infrastructure and consulting reviews.

How to Protect Yourself

  1. If you use any AI server or cache software, ask your provider or administrator whether LMCache is running, and if so whether it is set to listen only on a private or local address.
  2. Never expose an admin panel, cache service, or management port directly to the public internet; keep them behind a firewall or private network.
  3. Run server software with a limited user account instead of an administrator or root account so that a compromise cannot take over the whole machine.
  4. Apply patches and version updates as soon as vendors release them, and check the vendor's security page when a critical flaw is announced.
  5. If you cannot patch yet, reduce risk by restricting which computers can connect to the service, knowing that any connected computer could still attack it.
  6. Watch for an official fix from LMCache and do not rely only on third-party alerts for a workaround.

Vulnerabilities & Fixes

Terms Explained

  • LMCache Open source software that speeds up large language model servers by caching their repeated work.
  • unauthenticated Able to act without providing a username, password, or other proof of identity.
  • remote code execution A type of attack where an outsider can run their own commands on a computer.
  • pickle A Python data format that can carry code and run it when the data is read.
  • ZeroMQ A messaging library that lets different programs or machines pass messages to each other.
  • CVE A public identifier for a specific security vulnerability in software.
  • root The most powerful account on a Linux computer, which can change or access anything.
  • localhost The computer itself, not reachable from other machines.

Related AEU services

  • AEU-I IT and security consulting