CVE-2026-76850
CriticalCVSS 9.8Exploitation Probability (EPSS)
Elevated risk69th percentile - higher than 69% of all known CVEs
Summary
LMDeploy deserializes disaggregated-serving peer messages with pickle. The handle_zmq_recv coroutine in lmdeploy/pytorch/disagg/conn/engine_conn.py reads peer-to-peer cache-free requests with recv_pyobj(), which deserializes the received bytes with pickle.loads(), and the isinstance check against DistServeCacheFreeRequest runs only after deserialization has already completed. The peer that supplies those bytes is caller-controlled: p2p_connect passes remote_engine_endpoint_info.zmq_address from the request body to connect() on the ZMQ PULL socket, and the POST /distserve/p2p_initialize and /distserve/p2p_connect endpoints in lmdeploy/serve/openai/api_server.py apply no authentication unless the server is started with api_keys, which defaults to None. A remote attacker can direct an engine to pull from a ZMQ endpoint under their control and execute arbitrary code in the engine process. Deployments that do not enable disaggregated serving are not affected, because the receive loop is only started once the migration backend accepts the connection.
Risk Assessment
A remote attacker can execute arbitrary code in the engine process, potentially leading to full system compromise or data breach.
Recommendation
Update LMDeploy to a patched version, or enable authentication (api_keys) and restrict access to p2p endpoints.
Other vulnerabilities in LMDeploy
See all- CVE-2025-66455Critical
LMDeploy versions 0.9.2 through 0.16.0 used recv_pyobj() to deserialize messages received over a ZeroMQ PULL socket in the DistServe/PD-disaggregation control plane. PyZMQ implements recv_pyobj() with Python pickle deserialization, which can execute arbitrary code, and the peer address was supplied via the POST /distserve/p2p_connect HTTP endpoint. Without API-key authentication enabled, an attacker could cause the server to connect to an attacker-controlled ZeroMQ endpoint and deserialize a crafted pickle payload, leading to unauthenticated remote code execution.
- CVE-2025-59953Critical
LMDeploy implements an RPC server (AsyncRPCServer in zmq_rpc.py) whose call_and_response() function deserializes received messages using pickles.loads() without any sanitization. This allows remote code execution through the RPC server. The vulnerability exists from version 0.9.1 up to 0.10.2, where it is patched.
- CVE-2026-33625High
LMDeploy versions 0.12.1 through 0.12.2 contain a code injection vulnerability in lmdeploy/pytorch/config.py line 620. An attacker can execute arbitrary Python code by publishing a malicious HuggingFace model with a crafted quantization_config.quant_dtype value, which is passed to eval() without validation.
- CVE-2026-92983High
InternLM LMDeploy through 0.17.0 in DistServe prefill/decode disaggregation mode fails to release scheduler sessions because the proxy uses user-facing session IDs instead of internal scheduler keys. Unauthenticated attackers can send completion requests to the proxy endpoint that accumulate unreleased scheduler metadata and memory until the prefill worker is out-of-memory killed.
- CVE-2026-63764High
SSRF vulnerability in LMDeploy up to version 0.14.0 (fixed in commit 03c3130) in the _load_http_url function in connection.py. The private-IP guard only validates the original URL without re-validating hosts after HTTP redirects.
- CVE-2026-46517High
LMDeploy is a toolkit for compressing, deploying, and serving large language models. In versions 0.12.3 and prior, hardcoded "trust_remote_code=True" enables HF supply-chain RCE without user opt-in. Version 0.13.0 patches the issue.
- CVE-2026-46432High
LMDeploy versions up to 0.12.3 have hardcoded "trust_remote_code=True" in multiple HuggingFace model-loading call sites, allowing arbitrary code execution. No public patches are available at the time of publication.
Original NVD description (English source)
LMDeploy deserializes disaggregated-serving peer messages with pickle. The handle_zmq_recv coroutine in lmdeploy/pytorch/disagg/conn/engine_conn.py reads peer-to-peer cache-free requests with recv_pyobj(), which deserializes the received bytes with pickle.loads(), and the isinstance check against DistServeCacheFreeRequest runs only after deserialization has already completed. The peer that supplies those bytes is caller-controlled: p2p_connect passes remote_engine_endpoint_info.zmq_address from the request body to connect() on the ZMQ PULL socket, and the POST /distserve/p2p_initialize and /distserve/p2p_connect endpoints in lmdeploy/serve/openai/api_server.py apply no authentication unless the server is started with api_keys, which defaults to None. A remote attacker can direct an engine to pull from a ZMQ endpoint under their control and execute arbitrary code in the engine process. Deployments that do not enable disaggregated serving are not affected, because the receive loop is only started once the migration backend accepts the connection.
Vulnerability data from NVD (NIST) · CISA KEV · EPSS

