activity
20202024
most citedA Comprehensive Study of Deep Video Action Recognition

115 citations · 247 across the 15 of their papers we have counts for

collaborators

18 papers

cs.CV202216 cited

CoupAlign: Coupling Word-Pixel with Sentence-Mask Alignments for Referring Image Segmentation

Zicheng Zhang, Yi Zhu, Jianzhuang Liu +2

Referring image segmentation aims at localizing all pixels of the visual objects described by a natural language sentence. Previous works learn to straightforwardly align the sente…

cs.CV202222 cited

Visual Prompt Tuning for Test-time Domain Adaptation

Yunhe Gao, Xingjian Shi, Yi Zhu +5

Models should be able to adapt to unseen data during test-time to avoid performance drops caused by inevitable distribution shifts in real-world deployment scenarios. In this work,…

cs.CV20223 cited

ADAPT: Vision-Language Navigation with Modality-Aligned Action Prompts

Bingqian Lin, Yi Zhu, Zicong Chen +3

Vision-Language Navigation (VLN) is a challenging task that requires an embodied agent to perform action-level modality alignment, i.e., make instruction-asked actions sequentially…

cs.CV20221 cited

ImpDet: Exploring Implicit Fields for 3D Object Detection

Xuelin Qian, Li Wang, Yi Zhu +3

Conventional 3D object detection approaches concentrate on bounding boxes representation learning with several parameters, i.e., localization, dimension, and orientation. Despite i…

cs.CV20221 cited

BigDetection: A Large-scale Benchmark for Improved Object Detector Pre-training

Likun Cai, Zhi Zhang, Yi Zhu +3

Multiple datasets and open challenges for object detection have been introduced in recent years. To build more general and powerful object detection systems, in this paper, we cons…

cs.CV2021

Blending Anti-Aliasing into Vision Transformer

Shengju Qian, Hao Shao, Yi Zhu +2

The transformer architectures, based on self-attention mechanism and convolution-free design, recently found superior performance and booming applications in computer vision. Howev…