Cache keys that forget a dimension
Fixes in LiteLLM, TensorRT-LLM and vLLM for prompt and KV caches that silently stop hitting.
1,746 to 60cache-write tokens per reminder turn, before and after the LiteLLM fix
One bug shape, found in three serving stacks: a cache key that drops part of what makes a request distinct, or two components that derive the same key two different ways.
In LiteLLM, a mid-conversation system reminder on older Claude models was hoisted into the system prompt, which rewrote the cached prefix and re-billed the whole conversation at cache-write prices: about 1,746 written tokens on a turn that should write about 60. The fix converts the reminder in place. Two PRs merged on August 20, 2026, and a third was carried into #42630 with its authorship kept.
In TensorRT-LLM's KV-aware router, the LoRA id was left out of the routing key, so every LoRA request scored zero cache hits, and the router and the workers derived the cache salt with different hashes, so salted requests never matched a worker. That PR is open. In vLLM, a conformance suite for KV-cache key partitioning and a connector fix are open.