Skip to content
vLLMGHSA-83vm-p52w-f9pw

vLLM: extract_hidden_states speculative decoding crashes server on any request with penalty parameters

Medium6.5CVE-2026-44223 · Published May 6, 2026 · updated Sep 10, 2026

GitHub advisory

Affected versions

PackageAffectedFixed in
vllm
PyPI
>= 0.18.0, < 0.20.00.20.0
Details and references

### Summary The `extract_hidden_states` speculative decoding proposer in vLLM returns a tensor with an incorrect shape after the first decode step, causing a `RuntimeError` that crashes the EngineCore process. The crash is triggered when any request in the batch uses sampling penalty parameters (`repetition_penalty`, `frequency_penalty`, or `presence_penalty`). A single request with a penalty parameter (e.g., `"repetition_penalty": 1.1`) is sufficient to crash the server. The crash is deterministic and immediate , no concurrency, race condition, or special workload is required. ### Details In vLLM v0.17.0, the `extract_hidden_states` proposer's `propose()` method returned `sampled_token_ids.unsqueeze(-1)`, producing a tensor of shape `(batch_size, 1)`. In [PR #37013](https://github.com/vllm-project/vllm/pull/37013) (first released in v0.18.0), the KV connector interface was refactored out of `propose()`. The return type changed from `tuple[Tensor, KVConnectorOutput | None]` to `Tensor`, and the `.unsqueeze(-1)` call was removed along with the KV connector output: ```python # Before (v0.17.0): return sampled_token_ids.unsqueeze(-1), kv_connector_output # shape (batch_size, 1) # After (v0.18.0+): return sampled_token_ids # shape (batch_size, 2) after first decode step ``` The refactor missed that `sampled_token_ids` changed semantics between the first and subsequent decode steps. After the first decode step, the rejection sampler allocates its output as `(batch_size, max_spec_len + 1)`. With `num_speculative_tokens=1`, this produces shape `(batch_size, 2)` instead of the expected `(batch_size, 1)`, causing a broadcast shape mismatch during penalty application. ### Impact Any vLLM deployment between v0.18.0 and v0.19.1 (inclusive) configured with `extract_hidden_states` speculative decoding is affected. A single API request containing any penalty parameter immediately and permanently crashes the EngineCore process, resulting in complete loss of service availability. ### Patches Fixed in [PR #38610](https://github.com/vllm-project/vllm/pull/38610), first included in vLLM v0.20.0. The fix slices the return value to `sampled_token_ids[:, :1]`, ensuring the correct `(batch_size, 1)` shape regardless of the rejection sampler's output dimensions. ### Workarounds - Upgrade to vLLM v0.20.0 or later. - If upgrading is not possible, avoid using `extract_hidden_states` as the speculative decoding method on affected versions. - Alternatively, reject or strip penalty parameters (`repetition_penalty`, `frequency_penalty`, `presence_penalty`) from incoming requests at an API gateway before they reach vLLM.

CVSS 3.1
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
Severity from
GitHub (reviewed advisory)
Weakness
CWE-131, CWE-704
Also known as
CVE-2026-44223, PYSEC-2026-145

More vLLM advisories

All vLLM
DateAdvisory
May 5vLLM Vulnerable to Remote DoS via Special-Token Placeholders
CVE-2026-44222Medium6.5fixed in 0.20.0
Apr 27vLLM makes Use of Uninitialized Resource
CVE-2026-7141Low5.6fixed in 0.19.1
May 26vllm has Improper Resource Shutdown or Release
CVE-2026-9540Medium5.3no fix yet
Apr 3vLLM: Denial of Service via Unbounded Frame Count in video/jpeg Base64 Processing
CVE-2026-34755Medium6.5fixed in 0.19.0
Apr 3vLLM: Server-Side Request Forgery (SSRF) in `download_bytes_from_url `
CVE-2026-34753Medium5.4fixed in 0.19.0
Apr 3vLLM: Unauthenticated OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server
CVE-2026-34756Medium6.5fixed in 0.19.0

Critical advisories by email

Wednesdays: the week’s critical and high advisories in the AI and data stack, with the fixed versions. Only in weeks that have some.

Double opt-in. Unsubscribe any time.