vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion
Medium5.3CVE-2026-105760 · Published Oct 5, 2026 · updated Oct 6, 2026
## Summary The OpenAI-compatible chat endpoint accepts request-level video loader options through `media_io_kwargs`. A caller can select the GLMGA sampler and provide large `fps` and `max_frames` values. GLMGA first constructs an attacker-sized Python index list and then deduplicates it, even when the supplied video contains only a few frames. This work occurs in the shared media-loading executor before frame reads, allowing a tiny valid video and compact JSON options to consume disproportionate CPU time and memory and delay unrelated media requests. The demonstrated impact is partial denial of service. No memory corruption, data disclosure, or code execution is claimed. ## Affected configuration The server must expose chat completions for a video-capable model and accept request-level `media_io_kwargs`. The model does not need to use GLMGA by default because the request value overrides the model's loader mapping. The proof of concept uses an inline two-frame video and explicitly selects the OpenCV decoder, so no external media host or GPU decoder is required. The request is reachable without vLLM credentials when no API key is configured. With API-key authentication enabled, ...
Affected versions
| Package | Affected | Fixed in |
|---|---|---|
| vllm PyPI | >= 0.23.0rc2, < 0.30.0 | 0.30.0 |
Details and references
## Summary The OpenAI-compatible chat endpoint accepts request-level video loader options through `media_io_kwargs`. A caller can select the GLMGA sampler and provide large `fps` and `max_frames` values. GLMGA first constructs an attacker-sized Python index list and then deduplicates it, even when the supplied video contains only a few frames. This work occurs in the shared media-loading executor before frame reads, allowing a tiny valid video and compact JSON options to consume disproportionate CPU time and memory and delay unrelated media requests. The demonstrated impact is partial denial of service. No memory corruption, data disclosure, or code execution is claimed. ## Affected configuration The server must expose chat completions for a video-capable model and accept request-level `media_io_kwargs`. The model does not need to use GLMGA by default because the request value overrides the model's loader mapping. The proof of concept uses an inline two-frame video and explicitly selects the OpenCV decoder, so no external media host or GPU decoder is required. The request is reachable without vLLM credentials when no API key is configured. With API-key authentication enabled, any caller holding a valid key can reach the same path. ## Attack surface A remote caller submits a valid chat-completion request with a small video and the following request-level options: ```json { "media_io_kwargs": { "video": { "video_backend": "glmga", "backend": "opencv", "fps": 500000, "max_frames": 500000 } } } ``` Increasing both numeric values increases temporary list construction and deduplication work even when the final decoded frame set remains unchanged. ## Root cause 1. `ChatCompletionRequest` exposes `media_io_kwargs` as request-controlled nested values: [`protocol.py#L365-L371`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/entrypoints/openai/chat_completion/protocol.py#L365-L371). 2. The request options are carried into chat parameters without a numeric work bound on GLMGA's `fps` or `max_frames`: [`protocol.py#L571-L599`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/entrypoints/openai/chat_completion/protocol.py#L571-L599). 3. Media options are merged so request values override configured defaults, including `video_backend`: [`connector.py#L577-L605`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/multimodal/media/connector.py#L577-L605). 4. The connector selects a registered video loader and runs media loading through the shared executor: [`video.py#L28-L66`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/multimodal/media/video.py#L28-L66) and [`connector.py#L44-L47`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/multimodal/media/connector.py#L44-L47). 5. GLMGA calculates `extract_t = min(int(duration * fps), max_frames)`, builds a list with that many entries, and deduplicates it before decoding frames: [`video.py#L667-L740`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/multimodal/video.py#L667-L740). 6. The OpenCV decoder receives the already-deduplicated indices, so a two-frame file can trigger a large intermediate allocation while producing only two decoded frames: [`opencv.py#L28-L50`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/multimodal/video_decoders/opencv.py#L28-L50). The missing invariant is a strict upper bound on sampling work before the index list is constructed. Bounding decoded output after deduplication does not bound the vulnerable intermediate computation. ## Suggested remediation Validate request-level video sampling options before dispatching work to the media executor. Enforce conservative absolute limits for `fps`, `max_frames`, and especially the computed candidate count
- CVSS 3.1
- CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L
- Severity from
- GitHub (reviewed advisory)
- Weakness
- CWE-400
- Also known as
- CVE-2026-105760
More vLLM advisories
All vLLM| Date | Advisory | Severity | Fixed in |
|---|---|---|---|
| Oct 6 | vLLM: Harmony tool continuations drop `cache_salt` , restoring a cross-tenant prefix-cache membership oracle | Low3.1 | 0.30.0 |
| Oct 5 | vLLM: resource exhaustion | Medium5.3 | 0.30.0 |
| Oct 5 | vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features | Medium6.5 | 0.30.0 |
| Oct 5 | vLLM: improper input validation | Medium6.5 | 0.30.0 |
| Oct 5 | vLLM: insecure direct object reference | Medium4.2 | 0.30.0 |
| Oct 5 | vLLM: improper input validation | Medium6.5 | 0.30.0 |