Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs
This research paper explores the trade-offs between quality and serving cost when reusing a shared prefix key-value (KV) cache across LoRA adapters in a small-model deployment. The study finds that full-prefix reuse has a small quality difference and low prefill cost, while partial recomputation provides no demonstrated advantage.
Save an API key to vote.