1 paper
Kaitong Cai, Jusheng Zhang, Jing Yang +4
Large vision-language models (VLMs) typically process hundreds or thousands of visual tokens per image or video frame, incurring quadratic attention cost and substantial redundancy…