Skip to content
vLLMGHSA-mgrm-fgjv-mhv8

vLLM denial of service via outlines unbounded cache on disk

Medium6.5CVE-2025-29770 · Published Mar 19, 2025 · updated Aug 7, 2026

GitHub advisory

Affected versions

PackageAffectedFixed in
vllm
PyPI
< 0.8.00.8.0
Details and references

### Impact The [outlines](https://dottxt-ai.github.io/outlines/latest/) library is one of the backends used by vLLM to support structured output (a.k.a. guided decoding). Outlines provides an optional cache for its compiled grammars on the local filesystem. This cache has been on by default in vLLM. Outlines is also available by default through the OpenAI compatible API server. The affected code in vLLM is [vllm/model_executor/guided_decoding/outlines_logits_processors.py](https://github.com/vllm-project/vllm/blob/53be4a863486d02bd96a59c674bbec23eec508f6/vllm/model_executor/guided_decoding/outlines_logits_processors.py), which unconditionally uses the cache from outlines. vLLM should have this off by default and allow administrators to opt-in due to the potential for abuse. A malicious user can send a stream of very short decoding requests with unique schemas, resulting in an addition to the cache for each request. This can result in a Denial of Service if the filesystem runs out of space. Note that even if vLLM was configured to use a different backend by default, it is still possible to choose outlines on a per-request basis using the `guided_decoding_backend` key of the `extra_body` field of the request. This issue applies to the V0 engine only. The V1 engine is not affected. ### Patches * https://github.com/vllm-project/vllm/pull/14837 The fix is to disable this cache by default since it does not provide an option to limit its size. If you want to use this cache anyway, you may set the `VLLM_V0_USE_OUTLINES_CACHE` environment variable to `1`. ### Workarounds There is no way to workaround this issue in existing versions of vLLM other than preventing untrusted access to the OpenAI compatible API server. ### References

CVSS 3.1
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
Severity from
GitHub (reviewed advisory)
Weakness
CWE-770
Also known as
CVE-2025-29770, PYSEC-2025-223

More vLLM advisories

All vLLM
DateAdvisory
Mar 192025vLLM Allows Remote Code Execution via Mooncake Integration
CVE-2025-29783Critical9.0fixed in 0.8.0
Mar 202025vLLM Deserialization of Untrusted Data vulnerability
CVE-2024-11041Critical9.8no fix yet
Mar 202025vLLM allows Remote Code Execution by Pickle Deserialization via AsyncEngineRPCServer() RPC server entrypoints
CVE-2024-9053Critical9.8no fix yet
Mar 202025vLLM deserialization vulnerability in vllm.distributed.GroupCoordinator.recv_object
CVE-2024-9052Critical9.8no fix yet
Apr 152025vLLM vulnerable to Denial of Service by abusing xgrammar cache
GHSA-hf3c-wxg2-49q9Medium6.5fixed in 0.8.4
Apr 232025CVE-2025-24357 Malicious model remote code execution fix bypass with PyTorch < 2.6.0
GHSA-ggpf-24jw-3fcwCritical9.8fixed in 0.8.0

Critical advisories by email

Wednesdays: the week’s critical and high advisories in the AI and data stack, with the fixed versions. Only in weeks that have some.

Double opt-in. Unsubscribe any time.