AI and data stack advisories

Severe, 6 weeks2973Projects319

2973 severe, 6 weeks · 319 projects

vLLMPYSEC-2026-4184

vLLM through 0.29.0 fetches and fully materializes remote or inline media...

vLLM

CVE-2026-100650 · Published Sep 26, 2026 · updated Oct 7, 2026

Medium6.5
Fix: upgrade to 0.30.0 or later
Source advisory

vLLM through 0.29.0 fetches and fully materializes remote or inline media before enforcing its documented media controls (the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB compressed-audio size cap, default 25 MB, and the per-modality --limit-mm-per-prompt item limits). Across four ingress paths , the shared media-acquisition layer (HTTPConnection.get_bytes()/async_get_bytes()), the chat completions audio_url/base64 path, the batch speech runner, and the Rust frontend POST /tokenize route , the server reads the entire HTTP response body, base64-decodes the inline payload, or spawns one fetch/decode task per media part, and only then applies the limit (or, on some paths, never applies it). A remote attacker can therefore cause the API server or batch-runner process to allocate memory and consume outbound bandwidth proportional to an attacker-chosen body size or media item count before the request is rejected, resulting in pre-inference memory and bandwidth exhaustion (denial of service). The chat and batch surfaces require an API key when one is configured; the Rust frontend /tokenize route is unauthenticated by design. There is no code execution or data disclosure impact.

Affected versions

PackageAffectedFixed in
vllm
PyPI
< 0.30.00.30.0

Changes since it was listed

DateChange
Oct 8Severity: Unrated to Medium
Details and references

More vLLM advisories

All vLLM
Advisory
vLLM: denial of service
High7.5Sep 30
vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against...
High7.5Sep 26
vLLM: denial of service
High7.5Sep 21
vLLM: denial of service
High7.5Sep 21
vLLM: resource exhaustion
Medium5.3Sep 21
vLLM: attacker could allocate unbounded memory
High7.5Sep 21