vLLM: improper input validation
Medium6.5CVE-2026-105757 · Published Oct 5, 2026 · updated Oct 6, 2026
## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back each path reaches. ## Summary Three structured-output request paths let an ordinary request reach a condition that raises an **uncaught, engine-fatal** exception instead of a per-request validation error. The failure is not confined to the request that caused it: it escapes into `EngineCore`'s busy loop and triggers a fatal `_send_engine_dead()`, so one malformed structured-output request denies service to all concurrent and subsequent tenants of that engine. The requests are ordinary API calls; the client does not need any privileged configuration beyond the (often default) structured-output feature. The shared root cause is the absence of a per-request exception boundary around structured-output grammar/token handling: a value that should fail as a request-scoped validation error instead propagates as an uncaught exception (or bypasses frontend validation entirely) and reaches the schedul...
Affected versions
| Package | Affected | Fixed in |
|---|---|---|
| vllm PyPI | < 0.30.0 | 0.30.0 |
Details and references
## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back each path reaches. ## Summary Three structured-output request paths let an ordinary request reach a condition that raises an **uncaught, engine-fatal** exception instead of a per-request validation error. The failure is not confined to the request that caused it: it escapes into `EngineCore`'s busy loop and triggers a fatal `_send_engine_dead()`, so one malformed structured-output request denies service to all concurrent and subsequent tenants of that engine. The requests are ordinary API calls; the client does not need any privileged configuration beyond the (often default) structured-output feature. The shared root cause is the absence of a per-request exception boundary around structured-output grammar/token handling: a value that should fail as a request-scoped validation error instead propagates as an uncaught exception (or bypasses frontend validation entirely) and reaches the scheduler/engine loop, which treats the failure as fatal. These sites are distinct from the published fixes for [GHSA-6qc9-v4r8-22xg](https://github.com/vllm-project/vllm/security/advisories/GHSA-6qc9-v4r8-22xg) and [GHSA-8wr5-jm2h-8r4f](https://github.com/vllm-project/vllm/security/advisories/GHSA-8wr5-jm2h-8r4f): GHSA-6qc9 (PR #17623) added frontend json-schema/regex/type validation but the completion check on v0.25.1 still catches only `TimeoutError`, so Site 1's valid duplicate-root EBNF via the latched xgrammar backend still re-raises and kills EngineCore; GHSA-8wr5 (PR #44744) fixes recovered-token state in the eagle spec-decode path and explicitly does not treat trailing `-1` padding as the fault, so Site 2's `-1` reaching `backend_guidance.validate_tokens()` via `ngram_gpu` survives it. ## Affected code Links pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1). The three sites are: **Site 1 , latched-backend compile exception escapes.** `StructuredOutputManager` permanently latches one backend on the first structured-output request and always compiles through it, ignoring the per-request `auto` backend selection. A grammar that `auto` accepts only via a fallback backend (e.g. a duplicate-root EBNF) then makes the latched xgrammar compiler raise; the exception is stored in the grammar `Future`, and the completion check catches only `TimeoutError`, so `Future.result()` re-raises during scheduler promotion and kills EngineCore. - Backend defaults to `"auto"`: [`vllm/config/structured_outputs.py#L21`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/config/structured_outputs.py#L21). - The `auto` validator tries xgrammar and silently falls back, recording the resolved backend on the request: [`vllm/sampling_params.py#L1004-L1035`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/sampling_params.py#L1004-L1035) (fields at [`#L85-L88`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/sampling_params.py#L85-L88)). - The manager latches one backend and always compiles through it: [`vllm/v1/structured_output/__init__.py#L127-L159`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/structured_output/__init__.py#L127-L159) and `_create_grammar()` at [`#L173-L184`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/structured_output/__init__.py#L173-L184). - The xgrammar sink calls the native compiler unguarded: [`vllm/v1/structured_output/backend_xgrammar.py#L78-L110`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/structured_ou
More vLLM advisories
All vLLM| Date | Advisory | Severity | Fixed in |
|---|---|---|---|
| Oct 6 | vLLM: Harmony tool continuations drop `cache_salt` , restoring a cross-tenant prefix-cache membership oracle | Low3.1 | 0.30.0 |
| Oct 5 | vLLM: resource exhaustion | Medium5.3 | 0.30.0 |
| Oct 5 | vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion | Medium5.3 | 0.30.0 |
| Oct 5 | vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features | Medium6.5 | 0.30.0 |
| Oct 5 | vLLM: insecure direct object reference | Medium4.2 | 0.30.0 |
| Oct 5 | vLLM: improper input validation | Medium6.5 | 0.30.0 |