1 citations · 1 across the 4 of their papers we have counts for
5 papers
CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks
Mingshuang Luo, Ruibing Hou, Bo Chao +4
Human-centric visual analysis plays a pivotal role in diverse applications, including surveillance, healthcare, and human-computer interaction. With the emergence of large-scale un…
VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI
Chenqian Le, Yilin Zhao, Nikasadat Emami +4
Recent advances in fMRI-based visual decoding have enabled compelling reconstructions of perceived images. However, most approaches rely on subject-specific training, limiting scal…
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
Qiming Li, Zekai Ye, Xiaocheng Feng +8
Although Large Vision-Language Models (LVLMs) have demonstrated powerful capabilities in interpreting visual information, they frequently produce content that deviates from visual…
FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
Fan Yang, Yousong Zhu, Xin Li +6
Recent Large Vision Language Models (LVLMs) demonstrate promising capabilities in unifying visual understanding and generative modeling, enabling both accurate content understandin…
LSVOS Challenge Report: Large-scale Complex and Long Video Object Segmentation
Henghui Ding, Lingyi Hong, Chang Liu +30
Despite the promising performance of current video segmentation models on existing benchmarks, these models still struggle with complex scenes. In this paper, we introduce the 6th…