6 papers
CoCo-IR: Contextual Composed Image Retrieval
Shengcao Cao, Tanmaya Shekhar Dabral, Zhongli Ding +6
Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of complex, real-world visual search…
DCNNAnaCal: Physics-Informed Machine Learning for Accurate and Precise Weak Lensing Shear Estimation
Shurui Lin, Xiangchong Li, Ji Li +3
Traditional weak gravitational lensing shear estimators are carefully calibrated but struggle to fully capture realistic galaxy morphologies, point-spread-function (PSF) effects, b…
Refer to Any Segmentation Mask Group With Vision-Language Prompts
Shengcao Cao, Zijun Wei, Jason Kuen +6
Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for c…
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
Shengcao Cao, Liang-Yan Gui, Yu-Xiong Wang
Current large multimodal models (LMMs) face challenges in grounding, which requires the model to relate language components to visual entities. Contrary to the common practice that…
SOHES: Self-supervised Open-world Hierarchical Entity Segmentation
Shengcao Cao, Jiuxiang Gu, Jason Kuen +7
Open-world entity segmentation, as an emerging computer vision task, aims at segmenting entities in images without being restricted by pre-defined classes, offering impressive gene…
Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task Learning
Yuxiang Lu, Shengcao Cao, Yu-Xiong Wang
Vision Foundation Models (VFMs) have demonstrated outstanding performance on numerous downstream tasks. However, due to their inherent representation biases originating from differ…