1 paper · 1 filter
Jiahe Zhao, Rongkun Zheng, Yi Wang +2
In video Multimodal Large Language Models (video MLLMs), the visual encapsulation process plays a pivotal role in converting video contents into representative tokens for LLM input…