arXiv:2609.35908v1 Announce Type: new Abstract: Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity between the incoming query and cached queries. This design enables cache poisoning: an attacker can cache a malicious response under a query with high cosine similarity to benign requests.
Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning
About this summary. This is a short, independently written summary of an article first published by arXiv cs.CR. Cyber Security News did not report or verify the underlying story. Read the original: https://arxiv.org/abs/2609.35908
Source attribution: headline and facts are from arXiv cs.CR (arxiv.org). Summary method: excerpt of the source description. See our source attribution policy.





