vLLM: extract_hidden_states speculative decoding crashes server on any request with penalty parameters
Medium6.5CVE-2026-44223 · Published May 6, 2026 · updated Sep 10, 2026
Affected versions
| Package | Affected | Fixed in |
|---|---|---|
| vllm PyPI | >= 0.18.0, < 0.20.0 | 0.20.0 |
Details and references
### Summary The `extract_hidden_states` speculative decoding proposer in vLLM returns a tensor with an incorrect shape after the first decode step, causing a `RuntimeError` that crashes the EngineCore process. The crash is triggered when any request in the batch uses sampling penalty parameters (`repetition_penalty`, `frequency_penalty`, or `presence_penalty`). A single request with a penalty parameter (e.g., `"repetition_penalty": 1.1`) is sufficient to crash the server. The crash is deterministic and immediate , no concurrency, race condition, or special workload is required. ### Details In vLLM v0.17.0, the `extract_hidden_states` proposer's `propose()` method returned `sampled_token_ids.unsqueeze(-1)`, producing a tensor of shape `(batch_size, 1)`. In [PR #37013](https://github.com/vllm-project/vllm/pull/37013) (first released in v0.18.0), the KV connector interface was refactored out of `propose()`. The return type changed from `tuple[Tensor, KVConnectorOutput | None]` to `Tensor`, and the `.unsqueeze(-1)` call was removed along with the KV connector output: ```python # Before (v0.17.0): return sampled_token_ids.unsqueeze(-1), kv_connector_output # shape (batch_size, 1) # After (v0.18.0+): return sampled_token_ids # shape (batch_size, 2) after first decode step ``` The refactor missed that `sampled_token_ids` changed semantics between the first and subsequent decode steps. After the first decode step, the rejection sampler allocates its output as `(batch_size, max_spec_len + 1)`. With `num_speculative_tokens=1`, this produces shape `(batch_size, 2)` instead of the expected `(batch_size, 1)`, causing a broadcast shape mismatch during penalty application. ### Impact Any vLLM deployment between v0.18.0 and v0.19.1 (inclusive) configured with `extract_hidden_states` speculative decoding is affected. A single API request containing any penalty parameter immediately and permanently crashes the EngineCore process, resulting in complete loss of service availability. ### Patches Fixed in [PR #38610](https://github.com/vllm-project/vllm/pull/38610), first included in vLLM v0.20.0. The fix slices the return value to `sampled_token_ids[:, :1]`, ensuring the correct `(batch_size, 1)` shape regardless of the rejection sampler's output dimensions. ### Workarounds - Upgrade to vLLM v0.20.0 or later. - If upgrading is not possible, avoid using `extract_hidden_states` as the speculative decoding method on affected versions. - Alternatively, reject or strip penalty parameters (`repetition_penalty`, `frequency_penalty`, `presence_penalty`) from incoming requests at an API gateway before they reach vLLM.
More vLLM advisories
All vLLM| Date | Advisory | Severity | Fixed in |
|---|---|---|---|
| May 5 | vLLM Vulnerable to Remote DoS via Special-Token Placeholders CVE-2026-44222Medium6.5fixed in 0.20.0 | Medium6.5 | 0.20.0 |
| Apr 27 | vLLM makes Use of Uninitialized Resource CVE-2026-7141Low5.6fixed in 0.19.1 | Low5.6 | 0.19.1 |
| May 26 | vllm has Improper Resource Shutdown or Release CVE-2026-9540Medium5.3no fix yet | Medium5.3 | No fix yet |
| Apr 3 | vLLM: Denial of Service via Unbounded Frame Count in video/jpeg Base64 Processing CVE-2026-34755Medium6.5fixed in 0.19.0 | Medium6.5 | 0.19.0 |
| Apr 3 | vLLM: Server-Side Request Forgery (SSRF) in `download_bytes_from_url ` CVE-2026-34753Medium5.4fixed in 0.19.0 | Medium5.4 | 0.19.0 |
| Apr 3 | vLLM: Unauthenticated OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server CVE-2026-34756Medium6.5fixed in 0.19.0 | Medium6.5 | 0.19.0 |