Accelerating Dense LLMs via L0-regularized Mixture-of-Experts
A new method for accelerating dense Large Language Models (LLMs) using L0-regularized Mixture-of-Experts (MoE) approach is proposed. This method achieves a significant speedup of up to 2.5x without compromising performance, making it a promising solution for efficient LLM inference.
Save an API key to vote.