FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference
Researchers introduce FlexEE, an early exiting framework for large language model (LLM) inference that reduces computation and memory constraints in offloading-based deployments. FlexEE enables efficient early exit with minimal accuracy degradation, leading to significant speedups in LLM inference.
Save an API key to vote.