CVE-2026-54234
HighCVSS 7.5Exploitation Probability (EPSS)
Low risk26th percentile - higher than 26% of all known CVEs
Summary
In vLLM before version 0.24.0, a crafted speculative decoding request can cause the rejection sampler to produce a token outside the model vocabulary. This triggers a GPU device-side assertion crash in the engine worker, aborting all concurrent requests.
Risk Assessment
A remote attacker can exploit public gRPC Generate and Abort endpoints to crash the shared engine worker, causing denial of service (DoS) for all other clients until the worker is restarted.
Recommendation
Upgrade vLLM to version 0.24.0 or later immediately to mitigate this vulnerability.
Original NVD description (English source)
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, a frontend-legal multi-request speculative decoding workload can cause the rejection sampler to produce a recovered token equal to the model vocabulary size boundary value, which is then converted to negative one when the engine selects the next live token for a request and is written back into the drafter's input ids; that out-of-vocabulary value is later consumed by the model's embedding and attention path and crashes the engine worker with a GPU device-side assertion. The same triggering request sequence is reachable through the public gRPC Generate and Abort endpoints, so a remote client that can send generation requests can crash the shared engine worker, aborting concurrent requests and causing a service-wide denial of service for other clients of the deployment until the worker is restarted. This issue is fixed in version 0.24.0.

