1 paper
Chenghua Zhu, Zhaolu Kang, Qifan Shi +8
Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sam…