Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions
A study on optimal placement of KV caches across GPU, CPU, and SSD for long-lived sessions, with results showing tiering can support 73.02 times more concurrent sessions per GPU and lower cost per session by 62.04 times. The study compares placement policies such as recency, reuse frequency, predicted reuse, and EWMA predictor with prefetch lookahead across chat, agent, and document question answering workloads.
Save an API key to vote.