Scaling Audio Models Efficiently: Joint Optimization of Scale, Resolution, Adaptation, Precision, and Sparsity

Researchers present a framework for jointly optimizing large automatic speech recognition (ASR) models like Whisper across six dimensions: model size, temporal resolution, encoder token stride, low-rank adaptation capacity, weight precision, and sparsity pattern. This optimization aims to balance deployment objectives such as word error rate, inference FLOPs, and memory footprint.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.