1 citations · 1 across the 4 of their papers we have counts for
1 paper · 2 filters
Yecheng Wu, Zhuoyang Zhang, Junyu Chen +9
VILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for underst…