CVE Catalog

CVE-2026-105758

MediumCVSS 5.3
Published: Updated: Translated: NVD NIST

Exploitation Probability (EPSS)

Low risk
0.30%

21th percentile - higher than 21% of all known CVEs

Summary

In vLLM from 0.24.0 to 0.30.0, the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes accept request-level values for max_frames and fps without server-side ceilings, allowing an unauthenticated caller to decode every frame from attacker-controlled video and consume disproportionate memory.

Risk Assessment

An attacker can submit requests to /tokenize with large values, causing excessive frame decoding, memory consumption, and potential API process termination before scheduling.

Recommendation

Upgrade vLLM to version 0.30.0 or later, which includes a fix.

Other vulnerabilities in vLLM

See all
Original NVD description (English source)

vLLM is an inference and serving engine for large language models. From 0.24.0 until 0.30.0, the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes accept request-level values for the media_io_kwargs.video.max_frames and media_io_kwargs.video.fps fields without enforcing server-side ceilings. An unauthenticated caller can submit these values to the /tokenize endpoint, causing the sampler to decode every frame selected from attacker-controlled video input, consume disproportionate frontend memory, and potentially terminate the API process before scheduling or admission control. The Rust frontend is not affected because it rejects the media_io_kwargs field. This issue is fixed in version 0.30.0.

Vulnerability data from NVD (NIST) · CISA KEV · EPSS