most citedRegionBLIP: A Unified Multi-modal Pre-training Framework for Holistic and Regional Comprehension

4 citations · 10 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20232 cited

Viewpoint Integration and Registration with Vision Language Foundation Model for Image Change Understanding

Xiaonan Lu, Jianlong Yuan, Ruigang Niu +2

Recently, the development of pre-trained vision language foundation models (VLFMs) has led to remarkable performance in many tasks. However, these models tend to have strong single…

cs.CV20232 cited

ICPC: Instance-Conditioned Prompting with Contrastive Learning for Semantic Segmentation

Chaohui Yu, Qiang Zhou, Zhibin Wang +1

Modern supervised semantic segmentation methods are usually finetuned based on the supervised or self-supervised models pre-trained on ImageNet. Recent work shows that transferring…

cs.CV20234 cited

RegionBLIP: A Unified Multi-modal Pre-training Framework for Holistic and Regional Comprehension

Qiang Zhou, Chaohui Yu, Shaofeng Zhang +3

In this work, we investigate extending the comprehension of Multi-modal Large Language Models (MLLMs) to regional objects. To this end, we propose to extract features corresponding…

cs.CV2023

Improved Neural Radiance Fields Using Pseudo-depth and Fusion

Jingliang Li, Qiang Zhou, Chaohui Yu +4

Since the advent of Neural Radiance Fields, novel view synthesis has received tremendous attention. The existing approach for the generalization of radiance field reconstruction pr…

cs.CV20231 cited

Points-to-3D: Bridging the Gap between Sparse Points and Shape-Controllable Text-to-3D Generation

Chaohui Yu, Qiang Zhou, Jingliang Li +3

Text-to-3D generation has recently garnered significant attention, fueled by 2D diffusion models trained on billions of image-text pairs. Existing methods primarily rely on score d…