Pydantic AI: Concurrency-limited models can keep their slot when a streamed request ends early
> This issue was posted by Codex Desktop using gpt-6.1-sol on behalf of David.
Summary
Applications that wrap a model with ConcurrencyLimitedModel or limit_model_concurrency can permanently lose shared concurrency capacity when a streamed request releases its slot from a different task than the one that acquired it. This can happen when a stream ends early, and also when a stream is fully consumed using the default stream_text() debouncing.
In an application that exposes an affected streaming endpoint to network clients and shares a long-lived model limiter across requests, a client can repeatedly start a stream and disconnect. The completed requests retain their slots, eventually preventing subsequent requests that share the limiter from proceeding.
Agent-level max_concurrency and non-streaming model requests are not affected by this defect.
Details
The built-in limiter uses anyio.CapacityLimiter, which associates each acquired slot with its borrowing task. Pydantic AI's streaming lifecycle can acquire the slot on the task consuming the stream and run cleanup on another internal task. The limiter rejects that release, so the slot remains occupied even after the request has ended. Cleanup can raise a RuntimeError; a later request on the borrowing task can also fail because that task still holds a slot.
Early termination includes stopping iteration, a...
Affected versions
| Package | Affected | Fixed in |
|---|---|---|
| pydantic-ai PyPI | >= 2.10.0, < 2.53.0 | 2.53.0 |
Details and references
> This issue was posted by Codex Desktop using gpt-6.1-sol on behalf of David. ### Summary Applications that wrap a model with `ConcurrencyLimitedModel` or `limit_model_concurrency` can permanently lose shared concurrency capacity when a streamed request releases its slot from a different task than the one that acquired it. This can happen when a stream ends early, and also when a stream is fully consumed using the default `stream_text()` debouncing. In an application that exposes an affected streaming endpoint to network clients and shares a long-lived model limiter across requests, a client can repeatedly start a stream and disconnect. The completed requests retain their slots, eventually preventing subsequent requests that share the limiter from proceeding. Agent-level `max_concurrency` and non-streaming model requests are not affected by this defect. ### Details The built-in limiter uses `anyio.CapacityLimiter`, which associates each acquired slot with its borrowing task. Pydantic AI's streaming lifecycle can acquire the slot on the task consuming the stream and run cleanup on another internal task. The limiter rejects that release, so the slot remains occupied even after the request has ended. Cleanup can raise a `RuntimeError`; a later request on the borrowing task can also fail because that task still holds a slot. Early termination includes stopping iteration, a consumer exception, and cancellation. Fully consuming `stream_text()` with its default `debounce_by=0.1` can also reach the cross-task release path. Fully consumed streams must therefore not be assumed safe. ### Mitigation Upgrade to a patched release of `pydantic-ai` or `pydantic-ai-slim`. If you cannot upgrade yet, use the agent-level `max_concurrency` setting instead of a concurrency-limited model, or avoid streaming runs through a concurrency-limited model.
- CVSS 3.1
- CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H
- Severity from
- GitHub (reviewed advisory)
- Weakness
- CWE-772
- Also known as
- CVE-2026-107286
- github.com/pydantic/pydantic-ai/security/advisories/GHSA-6fqq-452j-qhrp
- nvd.nist.gov/vuln/detail/CVE-2026-107286
- github.com/pydantic/pydantic-ai/pull/9478
- github.com/pydantic/pydantic-ai/commit/453f19feeb7ab1d789f9393b1723c6a73b3d77b2
- github.com/pydantic/pydantic-ai
- github.com/pydantic/pydantic-ai/releases/tag/v2.53.0