CVE-2026-105753
MEDIUM
NVD
CVSS Score
6.5
Severity
MEDIUM
Published
Oct 05, 2026
Vendor
unknown
Description
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, the default mirrored multimodal LRU cache can commit a media hash in the frontend sender cache during multimodal rendering and before engine admission, while the engine receiver cache never receives the payload if that request is rejected. A later request reusing the same media hash causes MultiModalProcessorSenderCache to send no payload and MultiModalReceiverCache to reach an assertion with the message "Expected a cached item," producing a shared-service availability failure. This issue is fixed in version 0.28.0.
References
- https://github.com/vllm-project/vllm/commit/396204230423b7cc6798300926b8fa30190d26a9
- https://github.com/vllm-project/vllm/pull/46747
- https://github.com/vllm-project/vllm/pull/51897
- https://github.com/vllm-project/vllm/releases/tag/v0.28.0
- https://github.com/vllm-project/vllm/security/advisories/GHSA-ph3r-5jfg-f84f