검색 상세

CrossKV: A Bandwidth-Crossover-Aware Fallback Policy for Hierarchical KV Caches in LLM Serving

초록(요약문)

Large language model (LLM) serving systems reuse the key-value (KV) caches of recurring prefixes, such as system prompts, conversation histories, and document contexts, to reduce the prefill cost and the time-to-first-token (TTFT). To relieve the capacity limit of GPU memory, recent systems extend the KV cache to CPU DRAM and SSD tiers. In such hierarchical KV caches, however, a cache hit does not always reduce latency. This thesis empirically demonstrates a latency inversion in which a cold hit, where the KV resides only on the SSD, becomes slower than recomputing the prefix on the GPU, as measured on an A100 GPU with LMCache. Our analysis shows that the boundary of this inversion is determined by disk bandwidth, and that below a crossover bandwidth a cold hit is slower than recompute at every context length. Based on this observation, we propose CrossKV, a fallback policy that bypasses an unprofitable SSD cold hit by recomputing the prefix and promotes the recomputed KV to the CPU tier so that subsequent requests are served as warm hits. Across three workloads, CrossKV reduces the TTFT of reuse requests by up to 90%.

more

목차

Contents 4
List of Tables 6
List of Figures 7
Abstract 8
Chapter 1. Introduction 9
1.1 Background 9
1.2 Problem Statement: Latency Inversion in Hierarchical KV Caches 10
1.3 Research Content 10
1.4 Thesis Organization 11
Chapter 2. Background and Related Work 12
2.1 LLM Serving and the KV Cache 12
2.2 GPU-Internal Prefix KV Cache Reuse 13
2.3 Hierarchical KV Cache Offloading 14
2.4 KV Cache Selection and Compression 16
2.5 Limitations of Prior Work and the Position of This Thesis 16
Chapter 3. Analysis of Latency Characteristics of Hierarchical KV Caches 18
3.1 Classification of KV Access States 18
3.2 Measurement Methodology and TTFT per Access State 19
3.3 Latency Model and Crossover Point 20
3.4 Projection onto NVMe Environments 23
Chapter 4. Design and Implementation of CrossKV 25
4.1 Latency-Aware Fallback 25
4.2 Recompute-Based CPU-Tier Promotion and skip-store 26
4.3 Implementation 26
Chapter 5. Experiments and Evaluation 28
5.1 Experimental Setup 28
5.2 Fallback Effect per Workload 30
5.3 Throughput Analysis from a Disk-Bandwidth Perspective 31
5.4 Redundant-Write Controlled Experiment and skip-store 31
5.5 Conditional Value of the CPU-tier Promotion Policy 32
Chapter 6. Conclusion and Future Work 34
6.1 Conclusion 34
6.2 Limitations 34
6.3 Future Work 35
References 38
Appendix A. Raw Measurements per Access State 41

more