1 paper
Weitong Cai, Hang Zhang, Yukai Huang +6
Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text…