2 papers
cs.CV2025
Breaking the Encoder Barrier for Seamless Video-Language Understanding
Handong Li, Yiyuan Zhang, Longteng Guo +2
Most Video-Large Language Models (Video-LLMs) adopt an encoder-decoder framework, where a vision encoder extracts frame-wise features for processing by a language model. However, t…
cs.CV2024
Explore the Limits of Omni-modal Pretraining at Scale
Yiyuan Zhang, Handong Li, Jing Liu +1
We propose to build omni-modal intelligence, which is capable of understanding any modality and learning universal representations. In specific, we propose a scalable pretraining p…