Skip to content
vLLMGHSA-85xf-c7hm-whqw

vLLM: improper input validation

Medium6.5CVE-2026-105757 · Published Oct 5, 2026 · updated Oct 6, 2026

## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back each path reaches. ## Summary Three structured-output request paths let an ordinary request reach a condition that raises an **uncaught, engine-fatal** exception instead of a per-request validation error. The failure is not confined to the request that caused it: it escapes into `EngineCore`'s busy loop and triggers a fatal `_send_engine_dead()`, so one malformed structured-output request denies service to all concurrent and subsequent tenants of that engine. The requests are ordinary API calls; the client does not need any privileged configuration beyond the (often default) structured-output feature. The shared root cause is the absence of a per-request exception boundary around structured-output grammar/token handling: a value that should fail as a request-scoped validation error instead propagates as an uncaught exception (or bypasses frontend validation entirely) and reaches the schedul...

GitHub advisory

Affected versions

PackageAffectedFixed in
vllm
PyPI
< 0.30.00.30.0
Details and references

## Affected - **Ecosystem / package:** pip / `vllm` - **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back each path reaches. ## Summary Three structured-output request paths let an ordinary request reach a condition that raises an **uncaught, engine-fatal** exception instead of a per-request validation error. The failure is not confined to the request that caused it: it escapes into `EngineCore`'s busy loop and triggers a fatal `_send_engine_dead()`, so one malformed structured-output request denies service to all concurrent and subsequent tenants of that engine. The requests are ordinary API calls; the client does not need any privileged configuration beyond the (often default) structured-output feature. The shared root cause is the absence of a per-request exception boundary around structured-output grammar/token handling: a value that should fail as a request-scoped validation error instead propagates as an uncaught exception (or bypasses frontend validation entirely) and reaches the scheduler/engine loop, which treats the failure as fatal. These sites are distinct from the published fixes for [GHSA-6qc9-v4r8-22xg](https://github.com/vllm-project/vllm/security/advisories/GHSA-6qc9-v4r8-22xg) and [GHSA-8wr5-jm2h-8r4f](https://github.com/vllm-project/vllm/security/advisories/GHSA-8wr5-jm2h-8r4f): GHSA-6qc9 (PR #17623) added frontend json-schema/regex/type validation but the completion check on v0.25.1 still catches only `TimeoutError`, so Site 1's valid duplicate-root EBNF via the latched xgrammar backend still re-raises and kills EngineCore; GHSA-8wr5 (PR #44744) fixes recovered-token state in the eagle spec-decode path and explicitly does not treat trailing `-1` padding as the fault, so Site 2's `-1` reaching `backend_guidance.validate_tokens()` via `ngram_gpu` survives it. ## Affected code Links pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1). The three sites are: **Site 1 , latched-backend compile exception escapes.** `StructuredOutputManager` permanently latches one backend on the first structured-output request and always compiles through it, ignoring the per-request `auto` backend selection. A grammar that `auto` accepts only via a fallback backend (e.g. a duplicate-root EBNF) then makes the latched xgrammar compiler raise; the exception is stored in the grammar `Future`, and the completion check catches only `TimeoutError`, so `Future.result()` re-raises during scheduler promotion and kills EngineCore. - Backend defaults to `"auto"`: [`vllm/config/structured_outputs.py#L21`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/config/structured_outputs.py#L21). - The `auto` validator tries xgrammar and silently falls back, recording the resolved backend on the request: [`vllm/sampling_params.py#L1004-L1035`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/sampling_params.py#L1004-L1035) (fields at [`#L85-L88`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/sampling_params.py#L85-L88)). - The manager latches one backend and always compiles through it: [`vllm/v1/structured_output/__init__.py#L127-L159`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/structured_output/__init__.py#L127-L159) and `_create_grammar()` at [`#L173-L184`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/structured_output/__init__.py#L173-L184). - The xgrammar sink calls the native compiler unguarded: [`vllm/v1/structured_output/backend_xgrammar.py#L78-L110`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/structured_ou

CVSS 3.1
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
Severity from
GitHub (reviewed advisory)
Weakness
CWE-20, CWE-248, CWE-755
Also known as
CVE-2026-105757

More vLLM advisories

All vLLM

Critical advisories by email

Wednesdays: the week’s critical and high advisories in the AI and data stack, with the fixed versions. Only in weeks that have some.

Double opt-in. Unsubscribe any time.