Skip to content
vLLMGHSA-2phq-3phc-84px

vLLM: insecure direct object reference

Medium4.2CVE-2026-105755 · Published Oct 5, 2026 · updated Oct 6, 2026

## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the flash late-interaction query cache reaches. ## Summary On late-interaction `/score` and `/rerank` deployments with flash late interaction enabled (the default for supported models), the worker caches per-request query embeddings under a key derived from the **caller-controlled** `X-Request-Id` header. A second concurrent request that reuses the victim's header value replaces the victim's cached query embedding before document scoring , so the victim's documents are scored against the **attacker's** query. Because the data-parallel router pins all requests sharing a cache key to the same engine, the collision is deterministic for an attacker who reuses the victim's `X-Request-Id`. Depending on timing, one request can also consume the shared use counter and force the other request into a late-interaction cache-miss error. This is a remotely reachable, request-controlled cross-request int...

GitHub advisory

Affected versions

PackageAffectedFixed in
vllm
PyPI
< 0.30.00.30.0
Details and references

## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the flash late-interaction query cache reaches. ## Summary On late-interaction `/score` and `/rerank` deployments with flash late interaction enabled (the default for supported models), the worker caches per-request query embeddings under a key derived from the **caller-controlled** `X-Request-Id` header. A second concurrent request that reuses the victim's header value replaces the victim's cached query embedding before document scoring , so the victim's documents are scored against the **attacker's** query. Because the data-parallel router pins all requests sharing a cache key to the same engine, the collision is deterministic for an attacker who reuses the victim's `X-Request-Id`. Depending on timing, one request can also consume the shared use counter and force the other request into a late-interaction cache-miss error. This is a remotely reachable, request-controlled cross-request integrity break on the standard scoring and reranking endpoints. It requires only that flash late interaction be enabled, which is the default for supported models. ## Affected code Links pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1): - [`vllm/entrypoints/serve/engine/serving.py#L117-L124`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/serve/engine/serving.py#L117-L124) , `_base_request_id()` copies the public `X-Request-Id` header directly. - [`vllm/entrypoints/pooling/base/serving.py#L109`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/pooling/base/serving.py#L109) , the frontend request id is `f"{self.request_id_prefix}-{self._base_request_id(raw_request)}"`. - [`vllm/entrypoints/pooling/scoring/serving.py#L211`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/pooling/scoring/serving.py#L211) , `flash_late_interaction()` (at [L191](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/pooling/scoring/serving.py#L191)) derives worker cache keys directly from that id: `query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]`. - [`vllm/v1/pool/late_interaction.py#L30-L36`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/pool/late_interaction.py#L30-L36) , the data-parallel routing helper pins all requests sharing a `query_key` to the same engine via `crc32(query_key)`, making collisions deterministic. - [`vllm/v1/worker/gpu/pool/late_interaction_runner.py#L95`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/worker/gpu/pool/late_interaction_runner.py#L95) , the worker stores query embeddings in a process-local cache keyed only by that string: `self._query_cache[query_key] = output.clone()`. The caller-controlled header enters as the request id, and the flash late-interaction path derives the worker cache key directly from it: ```python # vllm/entrypoints/serve/engine/serving.py Lines 116-126 @staticmethod def _base_request_id( raw_request: Request | None, default: str | None = None ) -> str | None: """Pulls the request id to use from a header, if provided""" if raw_request is not None and ( (req_id := raw_request.headers.get("X-Request-Id")) is not None ): return req_id return random_uuid() if default is None else default ``` ```python # vllm/entrypoints/pooling/scoring/serving.py Lines 207-212 n_queries = ctx.n_queries n_docs = len(ctx.engine_inputs) - n_

CVSS 3.1
CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:N/I:L/A:L
Severity from
GitHub (reviewed advisory)
Weakness
CWE-639
Also known as
CVE-2026-105755

More vLLM advisories

All vLLM

Critical advisories by email

Wednesdays: the week’s critical and high advisories in the AI and data stack, with the fixed versions. Only in weeks that have some.

Double opt-in. Unsubscribe any time.