vLLM: OOM Denial of Service via Audio Decompression Bomb
Medium6.5CVE-2026-54233 · Published Jun 17, 2026 · updated Sep 10, 2026
### Summary vLLM's `/v1/audio/transcriptions` endpoint limits compressed upload size but not decoded PCM output. A 25MB OPUS file expands to ~14.9GB of float32 PCM at decode time. Tested on vLLM v0.19.0. ### Details `SpeechToTextProcessor` rejects uploads over `VLLM_MAX_AUDIO_CLIP_FILESIZE_MB` (default 25MB) based on compressed byte length, but the audio decoder in `audio.py` accumulates all decoded frames into memory with no size limit before returning: ```python # speech_to_text.py L184-189 if len(audio_data) / 1024 ** 2 > self.max_audio_filesize_mb: raise VLLMValidationError(...) y, sr = load_audio(buf, sr=self.asr_config.sample_rate) # decoded size unchecked # audio.py L77-107 chunks: list[npt.NDArray] = [] for frame in container.decode(stream): chunks.append(frame.to_ndarray()) audio = np.concatenate(chunks, axis=-1).astype(np.float32) # single contiguous allocation ``` A 25MB OPUS file at 6kbps encodes ~8.7 hours of audio. Decoding produces ~5.7GB of float32 PCM (232x amplification), and `np.concatenate` then allocates a second contiguous array, bringing peak RSS to ~14.9GB from a single request. `SpeechToTextConfig.max_audio_clip_s` (default 30s) applies only a...
Affected versions
| Package | Affected | Fixed in |
|---|---|---|
| vllm PyPI | < 0.24.0 | 0.24.0 |
Details and references
### Summary vLLM's `/v1/audio/transcriptions` endpoint limits compressed upload size but not decoded PCM output. A 25MB OPUS file expands to ~14.9GB of float32 PCM at decode time. Tested on vLLM v0.19.0. ### Details `SpeechToTextProcessor` rejects uploads over `VLLM_MAX_AUDIO_CLIP_FILESIZE_MB` (default 25MB) based on compressed byte length, but the audio decoder in `audio.py` accumulates all decoded frames into memory with no size limit before returning: ```python # speech_to_text.py L184-189 if len(audio_data) / 1024 ** 2 > self.max_audio_filesize_mb: raise VLLMValidationError(...) y, sr = load_audio(buf, sr=self.asr_config.sample_rate) # decoded size unchecked # audio.py L77-107 chunks: list[npt.NDArray] = [] for frame in container.decode(stream): chunks.append(frame.to_ndarray()) audio = np.concatenate(chunks, axis=-1).astype(np.float32) # single contiguous allocation ``` A 25MB OPUS file at 6kbps encodes ~8.7 hours of audio. Decoding produces ~5.7GB of float32 PCM (232x amplification), and `np.concatenate` then allocates a second contiguous array, bringing peak RSS to ~14.9GB from a single request. `SpeechToTextConfig.max_audio_clip_s` (default 30s) applies only after the full decode and does not prevent the allocation. ### Impact An unauthenticated attacker can exhaust server memory with a small number of concurrent requests, each a valid upload within the documented size limit. Severity was assessed with reference to prior OOM vulnerability reports in vLLM. ### Fix A fix for this vulnerability was merged here: https://github.com/vllm-project/vllm/pull/44970
- CVSS 3.1
- CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
- Severity from
- GitHub (reviewed advisory)
- Weakness
- CWE-409
- Also known as
- CVE-2026-54233, PYSEC-2026-3404
- github.com/vllm-project/vllm/security/advisories/GHSA-6pr9-rp53-2pmc
- nvd.nist.gov/vuln/detail/CVE-2026-54233
- github.com/vllm-project/vllm/pull/44970
- github.com/vllm-project/vllm/commit/1b1359c33269446f13c05da9a90c25174cbea590
- github.com/advisories/GHSA-6pr9-rp53-2pmc
- github.com/pypa/advisory-database/tree/main/vulns/vllm/PYSEC-2026-3404.yaml
- github.com/vllm-project/vllm
- github.com/vllm-project/vllm/releases/tag/v0.23.1rc0
- pypi.org/project/vllm
More vLLM advisories
All vLLM| Date | Advisory | Severity | Fixed in |
|---|---|---|---|
| Jun 17 | vLLM: incomplete CVE-2026-22778 fix leaks PIL repr addresses via Anthropic router | Medium5.3 | 0.24.0 |
| Jun 17 | vLLM: GGUF dequantize kernel int truncation exposes uninitialized GPU memory in multi-tenant serving | Medium7.5 | 0.24.0 |
| Jun 17 | ## Summary Issue 1: EXIF orientation not normalized → The image orientation... | Medium4.8 | 0.24.0 |
| Jun 17 | vLLM: temperature=NaN and temperature=Infinity bypass validation and propagate to GPU kernels | Medium6.5 | 0.24.0 |
| Jun 16 | vLLM: OpenAI auth bypass | Critical9.1 | 0.22.0 |
| Jun 16 | vLLM: remote code execution | High7.5 | 0.22.0 |