VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs
Researchers have introduced VideoMM, a framework that decouples selection from reasoning in video understanding tasks. It uses a macro proxy to select semantically relevant regions and projects them onto high-fidelity micro tokens for detailed understanding. This approach results in significant speedup and accuracy gains over current methods, making it a scalable paradigm for long-video understanding.
Save an API key to vote.