1 paper · 1 filter
Min Shi, Shihao Wang, Chieh-Yun Chen +6
Balancing temporal resolution and spatial detail under limited compute budget remains a key challenge for video-based multi-modal large language models (MLLMs). Existing methods ty…