Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Yecheng Wu, Zhuoyang Zhang, Junyu Chen +9
VILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for underst…
cs.CV2023
Semantic Complete Scene Forecasting from a 4D Dynamic Point Cloud Sequence
Zifan Wang, Zhuorui Ye, Haoran Wu +2
We study a new problem of semantic complete scene forecasting (SCSF) in this work. Given a 4D dynamic point cloud sequence, our goal is to forecast the complete scene corresponding…