vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features
Medium6.5CVE-2026-105754 · Published Oct 5, 2026 · updated Oct 6, 2026
## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the scale-out transport path reaches. ## Summary vLLM's disaggregated **scale-out** transport splits a multimodal request into a trusted render step (`POST /v1/chat/completions/render`) and a separate generate step (`POST /inference/v1/generate`). The generate route decodes a caller-supplied `features` object , serialized encoder tensors (`kwargs_data`), multimodal hashes (`mm_hashes`), placeholder ranges (`mm_placeholders`), and the internal field-processor selection , and forwards it into the engine **as if it had come from the trusted renderer**, with no rebinding to (or validation against) the active model's renderer contract. Because the two routes are ordinary auth-guarded HTTP endpoints (the `/inference` prefix is registered by default on generate-capable servers), any authenticated caller can submit an otherwise-valid render body with a single forged field. Depending on which fiel...
Affected versions
| Package | Affected | Fixed in |
|---|---|---|
| vllm PyPI | < 0.30.0 | 0.30.0 |
Details and references
## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the scale-out transport path reaches. ## Summary vLLM's disaggregated **scale-out** transport splits a multimodal request into a trusted render step (`POST /v1/chat/completions/render`) and a separate generate step (`POST /inference/v1/generate`). The generate route decodes a caller-supplied `features` object , serialized encoder tensors (`kwargs_data`), multimodal hashes (`mm_hashes`), placeholder ranges (`mm_placeholders`), and the internal field-processor selection , and forwards it into the engine **as if it had come from the trusted renderer**, with no rebinding to (or validation against) the active model's renderer contract. Because the two routes are ordinary auth-guarded HTTP endpoints (the `/inference` prefix is registered by default on generate-capable servers), any authenticated caller can submit an otherwise-valid render body with a single forged field. Depending on which field is forged, this produces: - an **engine-fatal crash** of the shared EngineCore process (denial of service), reproduced as a CUDA illegal-memory-access, a post-admission rank-mismatch `ValueError`, and a hard `assert` , three independent forged fields (sites 1, 2, 3); - **silent cross-request encoder-cache poisoning / disclosure** of shared encoder state when the cache-key hash is not bound to the payload (site 4); - **transport-level integrity loss** when sparse placeholder masks are dropped during render-to-generate replay (site 5). All five share one root cause and one fix shape: the reconstructed multimodal state on the scale-out path is trusted without being rebound to, and validated against, the active model's renderer output before it reaches the engine. These sites are distinct from prior multimodal hardening. Site 1 survives [GHSA-wv77-2vpf-vmmg](https://github.com/vllm-project/vllm/security/advisories/GHSA-wv77-2vpf-vmmg) (that fix validates full tensor shape in `MultiModalDataParser`/`get_input_embeddings` on the prompt-embeds path), because our request forges `image_grid_thw` metadata with the pixel bytes intact and reaches the Qwen2 vision RoPE/`cu_seqlens` and `image_embeds.split` sink, which the shape-check fix does not rebind. Site 4 is distinct from [GHSA-c65p-x677-fgj6](https://github.com/vllm-project/vllm/security/advisories/GHSA-c65p-x677-fgj6) (which folds metadata into `MultiModalHasher.serialize_item` to stop hash collisions), because the scale-out generate path trusts a caller-supplied `mm_hash` as the cache key with no origin binding, so that fix does not stop a caller from submitting a victim's hash or a `kwargs_data=None` cache read. ## Affected code Links pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1). Shared entry point and control surface for all five sites: - `POST /inference/v1/generate` route: [`vllm/entrypoints/scale_out/token_in_token_out/api_router.py#L46-L75`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/api_router.py#L46-L75). - Public `features` schema (`kwargs_data`, `mm_hashes`, `mm_placeholders`): [`vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L42-L63`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L42-L63). - `ServingTokens.serve_tokens()` copies the decoded geometry and hashes into engine structures without rebinding to the renderer schema: [`vllm/entrypoints/scale_out/token_in_token_out/serving.py#L145-L172`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_ou
More vLLM advisories
All vLLM| Date | Advisory | Severity | Fixed in |
|---|---|---|---|
| Oct 6 | vLLM: Harmony tool continuations drop `cache_salt` , restoring a cross-tenant prefix-cache membership oracle | Low3.1 | 0.30.0 |
| Oct 5 | vLLM: resource exhaustion | Medium5.3 | 0.30.0 |
| Oct 5 | vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion | Medium5.3 | 0.30.0 |
| Oct 5 | vLLM: improper input validation | Medium6.5 | 0.30.0 |
| Oct 5 | vLLM: insecure direct object reference | Medium4.2 | 0.30.0 |
| Oct 5 | vLLM: improper input validation | Medium6.5 | 0.30.0 |