1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2024
CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
Yeyuan Wang, Dehong Gao, Bin Li +7
The impressive performance of Large Language Model (LLM) has prompted researchers to develop Multi-modal LLM (MLLM), which has shown great potential for various multi-modal tasks.…
cs.CV2024★ 1 cited
Synthesizing Efficient Data with Diffusion Models for Person Re-Identification Pre-Training
Ke Niu, Haiyang Yu, Xuelin Qian +3
Existing person re-identification (Re-ID) methods principally deploy the ImageNet-1K dataset for model initialization, which inevitably results in sub-optimal situations due to the…
cs.CV2021
Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified Viewpoints
Jinyang Yuan, Bin Li, Xiangyang Xue
Visual scenes are extremely rich in diversity, not only because there are infinite combinations of objects and background, but also because the observations of the same scene may v…