1 paper
Zheng Wang, Haoran Chen, Haoxuan Qin +3
Long video understanding is challenging due to dense visual redundancy, long-range temporal dependencies, and the tendency of chain-of-thought and retrieval-based agents to accumul…