VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

Researchers have introduced VideoMM, a framework that decouples selection from reasoning in video understanding tasks. It uses a macro proxy to select semantically relevant regions and projects them onto high-fidelity micro tokens for detailed understanding. This approach results in significant speedup and accuracy gains over current methods, making it a scalable paradigm for long-video understanding.

RSS Score 0 9/16/2026, 4:00:00 AM Original Source
Save an API key to vote.