24 citations · 68 across the 7 of their papers we have counts for
7 papers
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
Guo Chen, Yifei Huang, Jilan Xu +7
Understanding videos is one of the fundamental directions in computer vision research, with extensive efforts dedicated to exploring various architectures such as RNN, 3D CNN, and…
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
Yuwen Xiong, Zhiqi Li, Yuntao Chen +10
We introduce Deformable Convolution v4 (DCNv4), a highly efficient and effective operator designed for a broad spectrum of vision applications. DCNv4 addresses the limitations of i…
Leveraging Vision-Centric Multi-Modal Expertise for 3D Object Detection
Linyan Huang, Zhiqi Li, Chonghao Sima +4
Current research is primarily dedicated to advancing the accuracy of camera-only 3D object detectors (apprentice) through the knowledge transferred from LiDAR- or multi-modal-based…
FB-BEV: BEV Representation from Forward-Backward View Transformations
Zhiqi Li, Zhiding Yu, Wenhai Wang +3
View Transformation Module (VTM), where transformations happen between multi-view image features and Bird-Eye-View (BEV) representation, is a crucial step in camera-based BEV perce…
PointGame: Geometrically and Adaptively Masked Auto-Encoder on Point Clouds
Yun Liu, Xuefeng Yan, Zhilei Chen +3
Self-supervised learning is attracting large attention in point cloud understanding. However, exploring discriminative and transferable features still remains challenging due to th…
RemoteTouch: Enhancing Immersive 3D Video Communication with Hand Touch
Yizhong Zhang, Zhiqi Li, Sicheng Xu +4
Recent research advance has significantly improved the visual realism of immersive 3D video communication. In this work we present a method to further enhance this immersive experi…