activity
20172023
most citedDPT: Deformable Patch-based Transformer for Visual Recognition

116 citations · 166 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV20232 cited

Mitigating Hallucination in Visual Language Models with Visual Supervision

Zhiyang Chen, Yousong Zhu, Yufei Zhan +4

Large vision-language models (LVLMs) suffer from hallucination a lot, generating responses that apparently contradict to the image content occasionally. The key problem lies in its…

cs.CV20228 cited

Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual Tasks

Zhiyang Chen, Yousong Zhu, Zhaowen Li +8

Visual tasks vary a lot in their output formats and concerned contents, therefore it is hard to process them with an identical structure. One main obstacle lies in the high-dimensi…

cs.CV2022

UniVIP: A Unified Framework for Self-Supervised Visual Pre-training

Zhaowen Li, Yousong Zhu, Fan Yang +9

Self-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images…

cs.CV2021116 cited

DPT: Deformable Patch-based Transformer for Visual Recognition

Zhiyang Chen, Yousong Zhu, Chaoyang Zhao +4

Transformer has achieved great success in computer vision, while how to split patches in an image remains a problem. Existing methods usually use a fixed-size patch embedding which…

cs.CV202036 cited

Identity-Guided Human Semantic Parsing for Person Re-Identification

Kuan Zhu, Haiyun Guo, Zhiwei Liu +2

Existing alignment-based methods have to employ the pretrained human parsing models to achieve the pixel-level alignment, and cannot identify the personal belongings (e.g., backpac…

cs.CV2019

Learning Feature Embeddings for Discriminant Model based Tracking

Linyu Zheng, Ming Tang, Yingying Chen +2

After observing that the features used in most online discriminatively trained trackers are not optimal, in this paper, we propose a novel and effective architecture to learn optim…