Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning
A new LLM decoding framework, Mahalanobis-Ensemble Decoding (ME-Decoding), is introduced. It frames candidate token selection as ensemble pruning, using Mahalanobis distance to enhance semantic diversity and preserve high probabilities. This approach offers a plug-and-play module with negligible inference overhead.
Save an API key to vote.