How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment
A research paper on optimizing key-value cache compression in reinforcement learning post-training to reduce memory usage, which is crucial for large language models.
Save an API key to vote.