CVE-2026-105758
MediumCVSS 5.3Exploitation Probability (EPSS)
Low risk21th percentile - higher than 21% of all known CVEs
Summary
In vLLM from 0.24.0 to 0.30.0, the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes accept request-level values for max_frames and fps without server-side ceilings, allowing an unauthenticated caller to decode every frame from attacker-controlled video and consume disproportionate memory.
Risk Assessment
An attacker can submit requests to /tokenize with large values, causing excessive frame decoding, memory consumption, and potential API process termination before scheduling.
Recommendation
Upgrade vLLM to version 0.30.0 or later, which includes a fix.
Other vulnerabilities in vLLM
See all- CVE-2026-48746Critical
Vulnerability in vLLM versions 0.3.0 to 0.22.0 allows bypass of OpenAI API AuthenticationMiddleware. An attacker can use the API without providing the configured VLLM_API_KEY or --api-key.
- CVE-2026-22778Critical
A vulnerability in vLLM from version 0.8.3 to 0.14.0 leaks a heap address when an invalid image is sent to the multimodal endpoint. This leak reduces ASLR effectiveness from 4 billion to about 8 guesses, facilitating further attacks.
- CVE-2026-105922Medium
A security flaw has been discovered in vllm-project vLLM up to 0.31.0, affecting the function get_token_bin_counts_and_mask in the file vllm/model_executor/layers/utils.py of the Penalty Handler component. Manipulation results in denial of service. Remote exploitation is possible, and the exploit has been publicly released.
- CVE-2026-105775Medium
A security vulnerability in vllm-project vLLM up to 0.31.0 causes an out-of-bounds read in the function conv_ssm_forward in the file mamba_mixer2.py, within the Completions Request Handler component. The attack can be carried out remotely, and the exploit has been publicly disclosed.
- CVE-2026-105760Medium
In vLLM before 0.30.0, a caller can use the media_io_kwargs field to select the GLMGA video backend and supply large fps and max_frames values without a strict work ceiling, causing disproportionate CPU and memory consumption.
- CVE-2026-105759Medium
In vLLM before 0.30.0, the Rust frontend's track_http_metrics middleware records the raw HTTP method token as a Prometheus label, allowing an unauthenticated attacker to send unique arbitrary method tokens and cause permanent creation of label sets, increasing memory usage.
- CVE-2026-105757Medium
In vLLM before 0.30.0, structured-output request failures can escape validation and reach the EngineCore fatal-error path, allowing ordinary constrained-generation requests to terminate the shared engine.
- CVE-2026-105756Medium
In vLLM before 0.30.0, OpenAI-compatible request models accept a non-empty cache_salt value without enforcing character and length restrictions, which can raise an uncaught ValueError during scheduler cache lookup, causing EngineCore to terminate.
- CVE-2026-105755Medium
In vLLM before 0.30.0, flash late-interaction scoring at /score and /rerank derives query_key from the caller-controlled X-Request-Id header, allowing a concurrent request to overwrite the cached query embedding of a victim.
- CVE-2026-105754Medium
In vLLM before 0.30.0, the /inference/v1/generate endpoint in the disaggregated scale-out path accepts caller-supplied tensors, cache identifiers, ranges, and wire-selected multimodal field processors without rebinding them to the active model renderer contract, which can terminate EngineCore, poison cache, or alter transport semantics.
Original NVD description (English source)
vLLM is an inference and serving engine for large language models. From 0.24.0 until 0.30.0, the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes accept request-level values for the media_io_kwargs.video.max_frames and media_io_kwargs.video.fps fields without enforcing server-side ceilings. An unauthenticated caller can submit these values to the /tokenize endpoint, causing the sampler to decode every frame selected from attacker-controlled video input, consume disproportionate frontend memory, and potentially terminate the API process before scheduling or admission control. The Rust frontend is not affected because it rejects the media_io_kwargs field. This issue is fixed in version 0.30.0.
Vulnerability data from NVD (NIST) · CISA KEV · EPSS

