CVE-2026-100650
MediumCVSS 6.5Exploitation Probability (EPSS)
Low risk47th percentile - higher than 47% of all known CVEs
Summary
vLLM through 0.29.0 fetches and fully materializes remote or inline media before enforcing its documented media controls (the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB compressed-audio size cap, default 25 MB, and per-modality --limit-mm-per-prompt item limits). Across four ingress paths the server reads the entire HTTP response body, base64-decodes the inline payload, or spawns one fetch/decode task per media part, and only then applies the limit (or never applies it). A remote attacker can cause pre-inference memory and bandwidth exhaustion (denial of service).
Risk Assessment
A remote attacker can exhaust memory and bandwidth of the API server or batch-runner process, causing a denial of service before inference is performed.
Recommendation
Upgrade vLLM to a version newer than 0.29.0 that enforces media limits before fetching, and secure the Rust frontend /tokenize route (unauthenticated by design).
Other vulnerabilities in vLLM
See all- CVE-2026-48746Critical
Vulnerability in vLLM versions 0.3.0 to 0.22.0 allows bypass of OpenAI API AuthenticationMiddleware. An attacker can use the API without providing the configured VLLM_API_KEY or --api-key.
- CVE-2026-22778Critical
A vulnerability in vLLM from version 0.8.3 to 0.14.0 leaks a heap address when an invalid image is sent to the multimodal endpoint. This leak reduces ASLR effectiveness from 4 billion to about 8 guesses, facilitating further attacks.
- CVE-2026-103241Medium
A flaw has been found in vllm-project vLLM up to 0.26.0. This vulnerability affects unknown code of the file rust/src/parser/src/unified/gemma4.rs of the component Gemma4UnifiedParser. Executing a manipulation can lead to denial of service. The attack may be launched remotely. The exploit has been published and may be used. Upgrading to version 0.29.1rc0 is able to resolve this issue. This patch is called 3439bad37e68ba9755a46f4f6b44a4aeaf1f60a9. Upgrading the affected component is advised.
- CVE-2026-100654Medium
vLLM before 0.29.0 accepts user-controlled stop_token_ids on the OpenAI-compatible POST /v1/completions and POST /v1/chat/completions endpoints but validates only that the values are integers, not that each token id is within the model vocabulary/logits range. When min_tokens > 0, the stop token ids are used as logits indices to suppress stop tokens, so an out-of-range id reaches a CUDA indexing operation (index_put_) and triggers a device-side assertion. An authenticated API user can send a single malformed completion request that returns 500 Internal Server Error and puts EngineCore into a fatal state, causing subsequent requests to fail until the service is restarted (denial of service).
- CVE-2026-100653Medium
In vLLM versions 0.22.1 through 0.28.0, the operator-supplied model revision pin (--revision / --code-revision) is not propagated to several Hugging Face artifact loads for the FunAudioChat and Tarsier2 architectures. This causes pinned deployments to still resolve behavior-affecting artifacts from the repository's default revision, potentially altering audio preprocessing, speech tokenizer behavior, or Tarsier2 configuration without any change to the operator's configured pin. This is a supply-chain integrity and reproducibility failure, but not remote code execution or a trust_remote_code=False bypass. Fixed in version 0.28.0.
- CVE-2026-100652Medium
vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends, allowing out-of-vocabulary token IDs to reach MinTokensLogitsProcessor. Attackers can submit requests with min_tokens greater than zero and out-of-vocabulary stop_token_ids to trigger CUDA tensor indexing failures that leave EngineCore in a fatal state requiring service restart.
- CVE-2026-100651Medium
vLLM before 0.29.0 fails to enforce decoder prompt-length validation on the disaggregated serving endpoint /inference/v1/generate. When the request contains a 'features' (multimodal) payload, caller-supplied token_ids are not checked against model_config.max_model_len, and for multimodal processors reporting skip_prompt_length_check=True an overlong prompt reaches the worker input-batch copy into a fixed max_model_len-wide NumPy row. A client able to reach the endpoint can trigger a worker failure and denial of service.
- CVE-2026-100649Low
vLLM before 0.29.0 contains a resource-limit bypass vulnerability in PyNvVideoCodec decoder allocation where sampler subclass shadowing allows independent counter increments. Unauthenticated attackers can select different sampler subclasses in video requests to exceed configured decoder limits and exhaust unaccounted GPU memory.
- CVE-2026-100648Medium
vLLM before 0.29.0 fails to enforce the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB limit in multimodal chat audio decoding, allowing unauthenticated clients to bypass file size restrictions. Attackers can submit oversized audio files through chat endpoints to consume excessive memory and CPU resources during decoding.
- CVE-2026-100647Medium
vLLM versions before 0.29.0 contain a denial-of-service vulnerability in the cache_salt parameter accepted on OpenAI-compatible and Anthropic API endpoints, which lacks maximum length validation and is processed on the single EngineCore scheduler thread. Unauthenticated attackers can send HTTP requests with multi-hundred-megabyte salt values that trigger expensive pickle serialization and SHA-256 hashing, stalling the scheduler thread and denying service to all concurrent requests.
Original NVD description (English source)
vLLM through 0.29.0 fetches and fully materializes remote or inline media before enforcing its documented media controls (the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB compressed-audio size cap, default 25 MB, and the per-modality --limit-mm-per-prompt item limits). Across four ingress paths — the shared media-acquisition layer (HTTPConnection.get_bytes()/async_get_bytes()), the chat completions audio_url/base64 path, the batch speech runner, and the Rust frontend POST /tokenize route — the server reads the entire HTTP response body, base64-decodes the inline payload, or spawns one fetch/decode task per media part, and only then applies the limit (or, on some paths, never applies it). A remote attacker can therefore cause the API server or batch-runner process to allocate memory and consume outbound bandwidth proportional to an attacker-chosen body size or media item count before the request is rejected, resulting in pre-inference memory and bandwidth exhaustion (denial of service). The chat and batch surfaces require an API key when one is configured; the Rust frontend /tokenize route is unauthenticated by design. There is no code execution or data disclosure impact.
Vulnerability data from NVD (NIST) · CISA KEV · EPSS

