3 citations · 6 across the 7 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024★ 2 cited
Vript: A Video Is Worth Thousands of Words
Dongjie Yang, Suyuan Huang, Chengqiang Lu +5
Advancements in multimodal learning, particularly in video understanding and generation, require high-quality video-text datasets for improved model performance. Vript addresses th…
cs.CV2024★ 3 cited
From Image to Video, what do we need in multimodal LLMs?
Suyuan Huang, Haoxin Zhang, Linqing Zhong +4
Covering from Image LLMs to the more complex Video LLMs, the Multimodal Large Language Models (MLLMs) have demonstrated profound capabilities in comprehending cross-modal informati…