vLLM: reachable assertion
Medium6.5CVE-2026-105753 · Published Oct 6, 2026
## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the mirrored sender/receiver cache protocol reaches. ## Summary vLLM's default multimodal cache (`mm_processor_cache_type="lru"`) mirrors state across two processes: the frontend (P0) holds only metadata (`MultiModalProcessorSenderCache`) while the engine core (P1) holds the real payload (`MultiModalReceiverCache`). The design invariant is that `get_and_update()` runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer "is this cached in P1?" without talking to P1. That invariant breaks when a request is **rejected after P0 has rendered and hashed the multimodal input** (populating the P0 cache) **but before P1 receives the item** , for example, an oversized chat prompt rejected on `max_model_len` *after* rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends `No...
Affected versions
| Package | Affected | Fixed in |
|---|---|---|
| vllm PyPI | < 0.28.0 | 0.28.0 |
Details and references
## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the mirrored sender/receiver cache protocol reaches. ## Summary vLLM's default multimodal cache (`mm_processor_cache_type="lru"`) mirrors state across two processes: the frontend (P0) holds only metadata (`MultiModalProcessorSenderCache`) while the engine core (P1) holds the real payload (`MultiModalReceiverCache`). The design invariant is that `get_and_update()` runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer "is this cached in P1?" without talking to P1. That invariant breaks when a request is **rejected after P0 has rendered and hashed the multimodal input** (populating the P0 cache) **but before P1 receives the item** , for example, an oversized chat prompt rejected on `max_model_len` *after* rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends `None` instead of the payload, and P1 , which has nothing cached , trips `assert mm_item is not None, f"Expected a cached item for {mm_hash=}"`. This is a remotely reachable, request-controlled cache-mirroring desync on the standard multimodal inference path. It requires only the default cache configuration. ## Affected code Links pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1): - **P0 metadata cache** , `MultiModalProcessorSenderCache` at [`vllm/multimodal/cache.py#L379`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L379); `get_and_update_item` at [`#L410-L421`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L410-L421), commit assertion at [`#L418`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L418). - **P1 payload cache (the actual sink)** , `MultiModalReceiverCache` at [`vllm/multimodal/cache.py#L630`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L630); `get_and_update_item` at [`#L652-L663`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L652-L663), with the failing `assert mm_item is not None, f"Expected a cached item for {mm_hash=}"` at [`#L660`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L660). - **Default `mm_processor_cache_type = "lru"`** at [`vllm/config/multimodal.py#L132`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/config/multimodal.py#L132); dispatch in [`vllm/multimodal/registry.py#L294-L307`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/registry.py#L294-L307) (sender) and [`#L322-L331`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/registry.py#L322-L331) (receiver). The `processor_only`, disabled-caching, and `shm` paths are not affected. - **Rejection-after-render window** , rendering happens before length validation in [`vllm/entrypoints/openai/chat_completion/serving.py#L206-L231`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/chat_completion/serving.py#L206-L231) (`render_chat_request`), and the length check raises after the render in [`vllm/entrypoints/serve/utils/api_utils.py#L171-L189`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/serve/utils/api_utils.py#L171-L189). - **Cache-commit call site** ,
- CVSS 3.1
- CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
- Severity from
- GitHub (reviewed advisory)
- Weakness
- CWE-617
- Also known as
- CVE-2026-105753
More vLLM advisories
All vLLM| Date | Advisory | Severity | Fixed in |
|---|---|---|---|
| Oct 6 | vLLM: Harmony tool continuations drop `cache_salt` , restoring a cross-tenant prefix-cache membership oracle | Low3.1 | 0.30.0 |
| Oct 5 | vLLM: resource exhaustion | Medium5.3 | 0.30.0 |
| Oct 5 | vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion | Medium5.3 | 0.30.0 |
| Oct 5 | vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features | Medium6.5 | 0.30.0 |
| Oct 5 | vLLM: improper input validation | Medium6.5 | 0.30.0 |
| Oct 5 | vLLM: insecure direct object reference | Medium4.2 | 0.30.0 |