Skip to content
vLLMGHSA-ph3r-5jfg-f84f

vLLM: reachable assertion

Medium6.5CVE-2026-105753 · Published Oct 6, 2026

## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the mirrored sender/receiver cache protocol reaches. ## Summary vLLM's default multimodal cache (`mm_processor_cache_type="lru"`) mirrors state across two processes: the frontend (P0) holds only metadata (`MultiModalProcessorSenderCache`) while the engine core (P1) holds the real payload (`MultiModalReceiverCache`). The design invariant is that `get_and_update()` runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer "is this cached in P1?" without talking to P1. That invariant breaks when a request is **rejected after P0 has rendered and hashed the multimodal input** (populating the P0 cache) **but before P1 receives the item** , for example, an oversized chat prompt rejected on `max_model_len` *after* rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends `No...

GitHub advisory

Affected versions

PackageAffectedFixed in
vllm
PyPI
< 0.28.00.28.0
Details and references

## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the mirrored sender/receiver cache protocol reaches. ## Summary vLLM's default multimodal cache (`mm_processor_cache_type="lru"`) mirrors state across two processes: the frontend (P0) holds only metadata (`MultiModalProcessorSenderCache`) while the engine core (P1) holds the real payload (`MultiModalReceiverCache`). The design invariant is that `get_and_update()` runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer "is this cached in P1?" without talking to P1. That invariant breaks when a request is **rejected after P0 has rendered and hashed the multimodal input** (populating the P0 cache) **but before P1 receives the item** , for example, an oversized chat prompt rejected on `max_model_len` *after* rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends `None` instead of the payload, and P1 , which has nothing cached , trips `assert mm_item is not None, f"Expected a cached item for {mm_hash=}"`. This is a remotely reachable, request-controlled cache-mirroring desync on the standard multimodal inference path. It requires only the default cache configuration. ## Affected code Links pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1): - **P0 metadata cache** , `MultiModalProcessorSenderCache` at [`vllm/multimodal/cache.py#L379`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L379); `get_and_update_item` at [`#L410-L421`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L410-L421), commit assertion at [`#L418`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L418). - **P1 payload cache (the actual sink)** , `MultiModalReceiverCache` at [`vllm/multimodal/cache.py#L630`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L630); `get_and_update_item` at [`#L652-L663`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L652-L663), with the failing `assert mm_item is not None, f"Expected a cached item for {mm_hash=}"` at [`#L660`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L660). - **Default `mm_processor_cache_type = "lru"`** at [`vllm/config/multimodal.py#L132`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/config/multimodal.py#L132); dispatch in [`vllm/multimodal/registry.py#L294-L307`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/registry.py#L294-L307) (sender) and [`#L322-L331`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/registry.py#L322-L331) (receiver). The `processor_only`, disabled-caching, and `shm` paths are not affected. - **Rejection-after-render window** , rendering happens before length validation in [`vllm/entrypoints/openai/chat_completion/serving.py#L206-L231`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/chat_completion/serving.py#L206-L231) (`render_chat_request`), and the length check raises after the render in [`vllm/entrypoints/serve/utils/api_utils.py#L171-L189`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/serve/utils/api_utils.py#L171-L189). - **Cache-commit call site** ,

CVSS 3.1
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
Severity from
GitHub (reviewed advisory)
Weakness
CWE-617
Also known as
CVE-2026-105753

More vLLM advisories

All vLLM

Critical advisories by email

Wednesdays: the week’s critical and high advisories in the AI and data stack, with the fixed versions. Only in weeks that have some.

Double opt-in. Unsubscribe any time.