4 papers
3D Question Answering via only 2D Vision-Language Models
Fengyun Wang, Sicheng Yu, Jiawei Wu +3
Large vision-language models (LVLMs) have significantly advanced numerous fields. In this work, we explore how to harness their potential to address 3D scene understanding tasks, u…
Mitigating Query Selection Bias in Referring Video Object Segmentation
Dingwei Zhang, Dong Zhang, Jinhui Tang
Recently, query-based methods have achieved remarkable performance in Referring Video Object Segmentation (RVOS) by using textual static object queries to drive cross-modal alignme…
Diffusion-Guided Knowledge Distillation for Weakly-Supervised Low-Light Semantic Segmentation
Chunyan Wang, Dong Zhang, Jinhui Tang
Weakly-supervised semantic segmentation aims to assign category labels to each pixel using weak annotations, significantly reducing manual annotation costs. Although existing metho…
R-Genie: Reasoning-Guided Generative Image Editing
Dong Zhang, Lingfeng He, Rui Yan +2
While recent advances in image editing have enabled impressive visual synthesis capabilities, current methods remain constrained by explicit textual instructions and limited editin…