RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
A new method called RBS-Attention is proposed to improve the efficiency of long-context large language models by reducing the cost of prefill. It uses two complementary selection branches to identify relevant tokens and achieves significant speedup and accuracy improvements.
Save an API key to vote.