1 paper · 1 filter
Xiao Yang, Yingzhe Ma, Haoxuan Yu +2
Long video understanding is heavily bottlenecked by a rigid one-shot paradigm: existing methods either densely encode videos at prohibitive memory and latency costs, or aggressivel…