116 citations · 132 across the 6 of their papers we have counts for
6 papers
Mitigating Hallucination in Visual Language Models with Visual Supervision
Zhiyang Chen, Yousong Zhu, Yufei Zhan +4
Large vision-language models (LVLMs) suffer from hallucination a lot, generating responses that apparently contradict to the image content occasionally. The key problem lies in its…
Masked Contrastive Pre-Training for Efficient Video-Text Retrieval
Fangxun Shu, Biaolong Chen, Yue Liao +6
We present a simple yet effective end-to-end Video-language Pre-training (VidLP) framework, Masked Contrastive Video-language Pretraining (MAC), for video-text retrieval tasks. Our…
Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual Tasks
Zhiyang Chen, Yousong Zhu, Zhaowen Li +8
Visual tasks vary a lot in their output formats and concerned contents, therefore it is hard to process them with an identical structure. One main obstacle lies in the high-dimensi…
UniVIP: A Unified Framework for Self-Supervised Visual Pre-training
Zhaowen Li, Yousong Zhu, Fan Yang +9
Self-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images…
DPT: Deformable Patch-based Transformer for Visual Recognition
Zhiyang Chen, Yousong Zhu, Chaoyang Zhao +4
Transformer has achieved great success in computer vision, while how to split patches in an image remains a problem. Existing methods usually use a fixed-size patch embedding which…
CoupleNet: Coupling Global Structure with Local Parts for Object Detection
Yousong Zhu, Chaoyang Zhao, Jinqiao Wang +3
The region-based Convolutional Neural Network (CNN) detectors such as Faster R-CNN or R-FCN have already shown promising results for object detection by combining the region propos…