CVE Catalog

CVE-2026-57173

MediumCVSS 6.5
Published: Updated: Translated: NVD NIST

Exploitation Probability (EPSS)

Elevated risk
0.69%

51th percentile - higher than 51% of all known CVEs

Summary

vLLM before version 0.24.0 in the input_audio handling path for /v1/chat/completions does not pass VLLM_MAX_AUDIO_DECODE_DURATION_S to the shared audio decoder. An unauthenticated client can submit a small compressed audio input that expands into a very large float32 PCM allocation, bypassing the duration guard and causing an out-of-memory worker crash.

Risk Assessment

An attacker can remotely cause a worker crash (denial of service) through memory exhaustion, affecting service availability. This affects deployments serving an audio-capable model.

Recommendation

Update vLLM to version 0.24.0 or later, which contains the fix for this issue.

Other vulnerabilities in vLLM

See all
Original NVD description (English source)

vLLM is an inference and serving engine for large language models. Prior to 0.24.0, the input_audio handling path for /v1/chat/completions calls AudioMediaIO.load_bytes or AudioMediaIO.load_file without passing VLLM_MAX_AUDIO_DECODE_DURATION_S to the shared audio decoder. An unauthenticated client can therefore submit a small compressed audio input that expands into a very large float32 PCM allocation, bypassing the duration guard already used by /v1/audio/transcriptions and causing an out-of-memory worker crash. Inline data URLs reach this path without being bounded by VLLM_AUDIO_FETCH_TIMEOUT. The issue affects deployments serving an audio-capable model, and authentication changes only the deployment-specific reachability. This issue is fixed in version 0.24.0.

Vulnerability data from NVD (NIST) · CISA KEV · EPSS