Skip to content
vLLMGHSA-58v5-2m8f-94pr

vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion

Medium5.3CVE-2026-105760 · Published Oct 5, 2026 · updated Oct 6, 2026

## Summary The OpenAI-compatible chat endpoint accepts request-level video loader options through `media_io_kwargs`. A caller can select the GLMGA sampler and provide large `fps` and `max_frames` values. GLMGA first constructs an attacker-sized Python index list and then deduplicates it, even when the supplied video contains only a few frames. This work occurs in the shared media-loading executor before frame reads, allowing a tiny valid video and compact JSON options to consume disproportionate CPU time and memory and delay unrelated media requests. The demonstrated impact is partial denial of service. No memory corruption, data disclosure, or code execution is claimed. ## Affected configuration The server must expose chat completions for a video-capable model and accept request-level `media_io_kwargs`. The model does not need to use GLMGA by default because the request value overrides the model's loader mapping. The proof of concept uses an inline two-frame video and explicitly selects the OpenCV decoder, so no external media host or GPU decoder is required. The request is reachable without vLLM credentials when no API key is configured. With API-key authentication enabled, ...

GitHub advisory

Affected versions

PackageAffectedFixed in
vllm
PyPI
>= 0.23.0rc2, < 0.30.00.30.0
Details and references

## Summary The OpenAI-compatible chat endpoint accepts request-level video loader options through `media_io_kwargs`. A caller can select the GLMGA sampler and provide large `fps` and `max_frames` values. GLMGA first constructs an attacker-sized Python index list and then deduplicates it, even when the supplied video contains only a few frames. This work occurs in the shared media-loading executor before frame reads, allowing a tiny valid video and compact JSON options to consume disproportionate CPU time and memory and delay unrelated media requests. The demonstrated impact is partial denial of service. No memory corruption, data disclosure, or code execution is claimed. ## Affected configuration The server must expose chat completions for a video-capable model and accept request-level `media_io_kwargs`. The model does not need to use GLMGA by default because the request value overrides the model's loader mapping. The proof of concept uses an inline two-frame video and explicitly selects the OpenCV decoder, so no external media host or GPU decoder is required. The request is reachable without vLLM credentials when no API key is configured. With API-key authentication enabled, any caller holding a valid key can reach the same path. ## Attack surface A remote caller submits a valid chat-completion request with a small video and the following request-level options: ```json { "media_io_kwargs": { "video": { "video_backend": "glmga", "backend": "opencv", "fps": 500000, "max_frames": 500000 } } } ``` Increasing both numeric values increases temporary list construction and deduplication work even when the final decoded frame set remains unchanged. ## Root cause 1. `ChatCompletionRequest` exposes `media_io_kwargs` as request-controlled nested values: [`protocol.py#L365-L371`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/entrypoints/openai/chat_completion/protocol.py#L365-L371). 2. The request options are carried into chat parameters without a numeric work bound on GLMGA's `fps` or `max_frames`: [`protocol.py#L571-L599`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/entrypoints/openai/chat_completion/protocol.py#L571-L599). 3. Media options are merged so request values override configured defaults, including `video_backend`: [`connector.py#L577-L605`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/multimodal/media/connector.py#L577-L605). 4. The connector selects a registered video loader and runs media loading through the shared executor: [`video.py#L28-L66`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/multimodal/media/video.py#L28-L66) and [`connector.py#L44-L47`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/multimodal/media/connector.py#L44-L47). 5. GLMGA calculates `extract_t = min(int(duration * fps), max_frames)`, builds a list with that many entries, and deduplicates it before decoding frames: [`video.py#L667-L740`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/multimodal/video.py#L667-L740). 6. The OpenCV decoder receives the already-deduplicated indices, so a two-frame file can trigger a large intermediate allocation while producing only two decoded frames: [`opencv.py#L28-L50`](https://github.com/vllm-project/vllm/blob/570735520942b6428ae98f1eb442c6d88adae55e/vllm/multimodal/video_decoders/opencv.py#L28-L50). The missing invariant is a strict upper bound on sampling work before the index list is constructed. Bounding decoded output after deduplication does not bound the vulnerable intermediate computation. ## Suggested remediation Validate request-level video sampling options before dispatching work to the media executor. Enforce conservative absolute limits for `fps`, `max_frames`, and especially the computed candidate count

CVSS 3.1
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L
Severity from
GitHub (reviewed advisory)
Weakness
CWE-400
Also known as
CVE-2026-105760

More vLLM advisories

All vLLM
Advisory
vLLM: Harmony tool continuations drop `cache_salt` , restoring a cross-tenant prefix-cache membership oracle
Low3.1Oct 6
vLLM: resource exhaustion
Medium5.3Oct 5
vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features
Medium6.5Oct 5
vLLM: improper input validation
Medium6.5Oct 5
vLLM: insecure direct object reference
Medium4.2Oct 5
vLLM: improper input validation
Medium6.5Oct 5

Critical advisories by email

Wednesdays: the week’s critical and high advisories in the AI and data stack, with the fixed versions. Only in weeks that have some.

Double opt-in. Unsubscribe any time.