Skip to content
vLLMPYSEC-2026-3998

vLLM before 0.29.0 validates allowed_token_ids against tokenizer length instead...

Medium5.3CVE-2026-93840 · Published Sep 18, 2026 · updated Sep 29, 2026

vLLM before 0.29.0 validates allowed_token_ids against tokenizer length instead of model output logits width in SamplingParams._validate_allowed_token_ids(). Attackers can supply token IDs above the output vocabulary that pass validation, causing LogitBiasState to corrupt GPU logits state and allow concurrent requests to sample tokens outside their allowlists.

Source advisory

Affected versions

PackageAffectedFixed in
vllm
PyPI
< 0.29.00.29.0
Details and references

More vLLM advisories

All vLLM

Critical advisories by email

Wednesdays: the week’s critical and high advisories in the AI and data stack, with the fixed versions. Only in weeks that have some.

Double opt-in. Unsubscribe any time.