1 paper
Xiao Yang, Yingzhe Ma, Haoxuan Yu +2
Long video understanding is heavily bottlenecked by a rigid one-shot paradigm: existing methods either densely encode videos at prohibitive memory and latency costs, or aggressivel…