1 paper
Lucy Lin, Ayush Jain, Yifan Liu +1
Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization an…