Attention-Aware Routing: Coupling Routing and Attention in MoEs
Researchers proposed Attention-Aware Routing (AAR), a new approach for routing in Mixture-of-Experts (MoE) language models. AAR uses temporal and spectral features from attention weights to improve model performance and reduce long diverging generation. The method exposes a retrieval-reasoning tension across depth in MoE models, making it a controlled probe for routing-relevant information.
Save an API key to vote.