13 citations · 37 across the 8 of their papers we have counts for
6 papers · 1 filter
SpatialBot: Precise Spatial Understanding with Vision Language Models
Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan +4
Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding, however they are still struggling with spatial understanding which is the foundation o…
FlexiFilm: Long Video Generation with Flexible Conditions
Yichen Ouyang, jianhao Yuan, Hao Zhao +2
Generating long and consistent videos has emerged as a significant yet challenging problem. While most existing diffusion-based video generation models, derived from image generati…
kNN-CLIP: Retrieval Enables Training-Free Segmentation on Continually Expanding Large Vocabularies
Zhongrui Gui, Shuyang Sun, Runjia Li +5
Continual segmentation has not yet tackled the challenge of improving open-vocabulary segmentation models with training data for accurate segmentation across large, continually exp…
SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model
Bin Cao, Jianhao Yuan, Yexin Liu +4
In the rapidly evolving area of image synthesis, a serious challenge is the presence of complex artifacts that compromise perceptual realism of synthetic images. To alleviate artif…
Efficient Multimodal Learning from Data-centric Perspective
Muyang He, Yexin Liu, Boya Wu +4
Multimodal Large Language Models (MLLMs) have demonstrated notable capabilities in general visual understanding and reasoning tasks. However, their deployment is hindered by substa…
Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
Jianhao Yuan, Francesco Pinto, Adam Davies +1
Neural image classifiers are known to undergo severe performance degradation when exposed to inputs that are sampled from environmental conditions that differ from their training d…