Skip to content
vLLMGHSA-x6mc-67gf-chw4

vLLM: resource exhaustion

Medium5.3CVE-2026-105758 · Published Oct 5, 2026 · updated Oct 6, 2026

### Summary An unauthenticated remote attacker can exhaust the memory of the vLLM API-server process by raising the request-level `media_io_kwargs.video.max_frames` and `fps` knobs on any deployment serving a Qwen2-VL or Qwen3-VL model. 74 extra bytes of JSON took the server's peak RSS from 2 271 MiB to 13 629 MiB over unauthenticated `POST /tokenize`. The `num_frames` ceiling reported in GHSA-vxqj-p4gw-9h4c and fixed by open PR #51969 does not reach this path: the Qwen samplers do not read `num_frames` at all. The same knobs were already capped upstream for `GLMGAVideoBackend` as an accepted security fix in `8b6de0eb9` (PR #54935, merged 2026-09-04); that cap never reached Qwen. ### Details #### Relationship to GHSA-vxqj-p4gw-9h4c and PR #51969 (read this first) GHSA-vxqj-p4gw-9h4c reported that request-level `media_io_kwargs.video.num_frames` overrides the engine frame-count ceiling, and open PR #51969 fixes it by clamping `num_frames` inside `VideoMediaIO.merge_kwargs`. That clamp does not reach the Qwen samplers. `Qwen2VLVideoBackend` and `Qwen3VLVideoBackend` do not read `num_frames` at all , `Qwen2VLVideoBackend`'s own docstring says so ("``num_frames`` is ignored (fps...

GitHub advisory

Affected versions

PackageAffectedFixed in
vllm
PyPI
>= 0.24.0, < 0.30.00.30.0
Details and references

### Summary An unauthenticated remote attacker can exhaust the memory of the vLLM API-server process by raising the request-level `media_io_kwargs.video.max_frames` and `fps` knobs on any deployment serving a Qwen2-VL or Qwen3-VL model. 74 extra bytes of JSON took the server's peak RSS from 2 271 MiB to 13 629 MiB over unauthenticated `POST /tokenize`. The `num_frames` ceiling reported in GHSA-vxqj-p4gw-9h4c and fixed by open PR #51969 does not reach this path: the Qwen samplers do not read `num_frames` at all. The same knobs were already capped upstream for `GLMGAVideoBackend` as an accepted security fix in `8b6de0eb9` (PR #54935, merged 2026-09-04); that cap never reached Qwen. ### Details #### Relationship to GHSA-vxqj-p4gw-9h4c and PR #51969 (read this first) GHSA-vxqj-p4gw-9h4c reported that request-level `media_io_kwargs.video.num_frames` overrides the engine frame-count ceiling, and open PR #51969 fixes it by clamping `num_frames` inside `VideoMediaIO.merge_kwargs`. That clamp does not reach the Qwen samplers. `Qwen2VLVideoBackend` and `Qwen3VLVideoBackend` do not read `num_frames` at all , `Qwen2VLVideoBackend`'s own docstring says so ("``num_frames`` is ignored (fps-driven, like the Qwen3-VL loader)"). They bound on `max_frames`, read from the same merged dict with no ceiling: ```python # vllm/multimodal/video.py , Qwen3VLVideoBackend.compute_frames_index_to_sample min_frames = kwargs.get("min_frames", 4) max_frames = kwargs.get("max_frames", 768) num_frames = int(total_frames_num / original_fps * fps) num_frames = min(max(num_frames, min_frames), max_frames, total_frames_num) ``` With `max_frames` raised from the request, the only remaining bound is `total_frames_num` , every frame in the container. I applied PR #51969's patch locally and re-ran both paths through the real merge layer (`merge_media_io_kwargs` → `VideoMediaIO.merge_kwargs` → `MediaConnector`), against `main` @ `b23433088`: | request `media_io_kwargs.video` | merged kwargs after #51969 | frames decoded | peak RSS | |---|---|---|---| | *(absent)* | `None` | 32 | 706 MiB | | `{"video_backend":"opencv","num_frames":-1}` | `{…,"num_frames":32}` | **32 , fixed** | 706 MiB | | `{"video_backend":"qwen3_vl"}` | `{…,"num_frames":32}` | 60 | 779 MiB | | `{"video_backend":"qwen3_vl","max_frames":1e9,"fps":1e6}` | `{…,"max_frames":1000000000,"fps":1000000,"num_frames":32}` | **900 , survives** | **2 994 MiB** | #51969 does exactly what it claims for `num_frames`; the clamp writes `num_frames: 32` into the merged dict and the Qwen sampler ignores it, while `max_frames` and `fps` pass through untouched. #### The codebase already has the fix pattern, on other backends This is not a new control being proposed. Commit `8b6de0eb9` , "[Security] Cap GLMGA video sampling to prevent request-driven resource exhaustion" (PR #54935, merged 2026-09-04, same author as #51969) , caps precisely these two knobs for `GLMGAVideoBackend`: ```python _MAX_FRAMES: ClassVar[int] = 640 _MAX_FPS: ClassVar[int] = 30 ... target_fps = min(target.fps, cls._MAX_FPS) max_frames = min(kwargs.get("max_frames", cls._MAX_FRAMES), cls._MAX_FRAMES) ``` Its description states the root cause as "both `target_fps` and `max_frames` are controllable via request-level `media_io_kwargs`", and says it follows "the pattern established by `GLM46VVideoBackend`" (which caps via `_MAX_FRAME_COUNT_DYNAMIC = 640` and `_MAX_DURATION = 2400`). So two backends cap request-controlled `fps`/`max_frames` as an accepted security measure. `Qwen2VLVideoBackend` and `Qwen3VLVideoBackend` , the most widely deployed video models on vLLM , cap neither. In GLMGA the uncapped knobs sized an intermediate *index list*; in the Qwen samplers they size the *decoded frame buffer*, which is larger by the per-frame pixel count. #### Reachability , default configuration, no authentication * `media_io_kwargs` is a request body field on `ChatCompletionRequest` and is carried by `/v1/chat/completions`, `/v1/embeddi

CVSS 3.1
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L
Severity from
GitHub (reviewed advisory)
Weakness
CWE-770
Also known as
CVE-2026-105758

More vLLM advisories

All vLLM

Critical advisories by email

Wednesdays: the week’s critical and high advisories in the AI and data stack, with the fixed versions. Only in weeks that have some.

Double opt-in. Unsubscribe any time.