CVE-2026-92983
HighCVSS 7.5Exploitation Probability (EPSS)
Low risk30th percentile - higher than 30% of all known CVEs
Summary
InternLM LMDeploy through 0.17.0 in DistServe prefill/decode disaggregation mode fails to release scheduler sessions because the proxy uses user-facing session IDs instead of internal scheduler keys. Unauthenticated attackers can send completion requests to the proxy endpoint that accumulate unreleased scheduler metadata and memory until the prefill worker is out-of-memory killed.
Risk Assessment
An attacker can cause memory exhaustion and shutdown of the model serving service, resulting in system unavailability (DoS).
Recommendation
Update LMDeploy to a version later than 0.17.0 and restrict access to the proxy endpoint to trusted clients.
Other vulnerabilities in LMDeploy
See all- CVE-2026-33625High
LMDeploy versions 0.12.1 through 0.12.2 contain a code injection vulnerability in lmdeploy/pytorch/config.py line 620. An attacker can execute arbitrary Python code by publishing a malicious HuggingFace model with a crafted quantization_config.quant_dtype value, which is passed to eval() without validation.
- CVE-2025-66455Critical
LMDeploy versions 0.9.2 through 0.16.0 used recv_pyobj() to deserialize messages received over a ZeroMQ PULL socket in the DistServe/PD-disaggregation control plane. PyZMQ implements recv_pyobj() with Python pickle deserialization, which can execute arbitrary code, and the peer address was supplied via the POST /distserve/p2p_connect HTTP endpoint. Without API-key authentication enabled, an attacker could cause the server to connect to an attacker-controlled ZeroMQ endpoint and deserialize a crafted pickle payload, leading to unauthenticated remote code execution.
- CVE-2025-59953Critical
LMDeploy implements an RPC server (AsyncRPCServer in zmq_rpc.py) whose call_and_response() function deserializes received messages using pickles.loads() without any sanitization. This allows remote code execution through the RPC server. The vulnerability exists from version 0.9.1 up to 0.10.2, where it is patched.
- CVE-2026-76850Critical
LMDeploy deserializes disaggregated-serving peer messages with pickle. The handle_zmq_recv coroutine in lmdeploy/pytorch/disagg/conn/engine_conn.py reads peer-to-peer cache-free requests with recv_pyobj(), which deserializes the received bytes with pickle.loads(), and the isinstance check against DistServeCacheFreeRequest runs only after deserialization has already completed. The peer that supplies those bytes is caller-controlled: p2p_connect passes remote_engine_endpoint_info.zmq_address from the request body to connect() on the ZMQ PULL socket, and the POST /distserve/p2p_initialize and /distserve/p2p_connect endpoints in lmdeploy/serve/openai/api_server.py apply no authentication unless the server is started with api_keys, which defaults to None. A remote attacker can direct an engine to pull from a ZMQ endpoint under their control and execute arbitrary code in the engine process. Deployments that do not enable disaggregated serving are not affected, because the receive loop is only started once the migration backend accepts the connection.
- CVE-2026-63764High
SSRF vulnerability in LMDeploy up to version 0.14.0 (fixed in commit 03c3130) in the _load_http_url function in connection.py. The private-IP guard only validates the original URL without re-validating hosts after HTTP redirects.
- CVE-2026-46517High
LMDeploy is a toolkit for compressing, deploying, and serving large language models. In versions 0.12.3 and prior, hardcoded "trust_remote_code=True" enables HF supply-chain RCE without user opt-in. Version 0.13.0 patches the issue.
- CVE-2026-46432High
LMDeploy versions up to 0.12.3 have hardcoded "trust_remote_code=True" in multiple HuggingFace model-loading call sites, allowing arbitrary code execution. No public patches are available at the time of publication.
Original NVD description (English source)
InternLM LMDeploy through 0.17.0 in DistServe prefill/decode disaggregation mode fails to release scheduler sessions because the proxy uses user-facing session IDs instead of internal scheduler keys. Unauthenticated attackers can send completion requests to the proxy endpoint that accumulate unreleased scheduler metadata and memory until the prefill worker is out-of-memory killed.

