1 paper · 1 filter
Yu Chen, Xiaohong Li, Xiaole Wang +3
In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply.…