1 paper · 1 filter
Jiazheng Li, Chi-Hao Wu, Yunze Liu +3
Understanding ultra-long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For current multimodal LLMs: even…