Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Towards Semantic Equivalence of Tokenization in Multimodal LLM
Shengqiong Wu, Hao Fei, Xiangtai Li +4
Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in processing vision-language tasks. One of the crux of MLLMs lies in vision tokenization, which…
cs.CV2024
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
Hao Fei, Shengqiong Wu, Meishan Zhang +3
While pre-training large-scale video-language models (VLMs) has shown remarkable potential for various downstream video-language tasks, existing VLMs can still suffer from certain…