2 papers
cs.CV2024
SpatialBot: Precise Spatial Understanding with Vision Language Models
Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan +4
Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding, however they are still struggling with spatial understanding which is the foundation o…
cs.CV2024
FlexiFilm: Long Video Generation with Flexible Conditions
Yichen Ouyang, jianhao Yuan, Hao Zhao +2
Generating long and consistent videos has emerged as a significant yet challenging problem. While most existing diffusion-based video generation models, derived from image generati…