Publications (56)
Learnable Tree Filter for Structure-preserving Feature Transform
Lin Song, Yanwei Li, Zeming Li +4
Learning discriminative global features plays a vital role in semantic segmentation. And most of the existing methods adopt stacks of local convolutions or non-local blocks to capt…
Voxel Field Fusion for 3D Object Detection
Yanwei Li, Xiaojuan Qi, Yukang Chen +4
In this work, we present a conceptually simple yet effective framework for cross-modality 3D object detection, named voxel field fusion. The proposed approach aims to maintain cros…
HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors
Panwang Pan, Zhuo Su, Chenguo Lin +6
Despite recent advancements in high-fidelity human reconstruction techniques, the requirements for densely captured images or time-consuming per-instance optimization significantly…
Unifying Voxel-based Representation with Transformer for 3D Object Detection
Yanwei Li, Yilun Chen, Xiaojuan Qi +3
In this work, we present a unified framework for multi-modality 3D object detection, named UVTR. The proposed method aims to unify multi-modality representations in the voxel space…
DetNet: A Backbone network for Object Detection
Zeming Li, Chao Peng, Gang Yu +3
Recent CNN based object detectors, no matter one-stage methods like YOLO, SSD, and RetinaNe or two-stage detectors like Faster R-CNN, R-FCN and FPN are usually trying to directly f…
IQDet: Instance-wise Quality Distribution Sampling for Object Detection
Yuchen Ma, Songtao Liu, Zeming Li +1
We propose a dense object detector with an instance-wise sampling strategy, named IQDet. Instead of using human prior sampling strategies, we first extract the regional feature of…
Rethinking Learnable Tree Filter for Generic Feature Transform
Lin Song, Yanwei Li, Zhengkai Jiang +5
The Learnable Tree Filter presents a remarkable approach to model structure-preserving relations for semantic segmentation. Nevertheless, the intrinsic geometric constraint forces…
Real-time Object Detection for Streaming Perception
Jinrong Yang, Songtao Liu, Zeming Li +2
Autonomous driving requires the model to perceive the environment and (re)act within a low latency for safety. While past works ignore the inevitable changes in the environment aft…
UREM: A High-performance Unified and Resilient Enhancement Method for Multi- and High-Dimensional Indexes
Ming Sheng, Shuliang Wang, Yong Zhang +3
Numerous multi- or high-dimensional indexes with distinct advantages have been proposed on various platforms to meet application requirements. To achieve higher-performance queries…
Dynamic Grained Encoder for Vision Transformers
Lin Song, Songyang Zhang, Songtao Liu +5
Transformers, the de-facto standard for language modeling, have been recently applied for vision tasks. This paper introduces sparse queries for vision transformers to exploit the…
4K4DGen: Panoramic 4D Generation at 4K Resolution
Renjie Li, Panwang Pan, Bangbang Yang +8
The blooming of virtual reality and augmented reality (VR/AR) technologies has driven an increasing demand for the creation of high-quality, immersive, and dynamic environments. Ho…
BEVStereo: Enhancing Depth Estimation in Multi-view 3D Object Detection with Dynamic Temporal Stereo
Yinhao Li, Han Bao, Zheng Ge +3
Bounded by the inherent ambiguity of depth perception, contemporary camera-based 3D object detection methods fall into the performance bottleneck. Intuitively, leveraging temporal…
PersDet: Monocular 3D Detection in Perspective Bird's-Eye-View
Hongyu Zhou, Zheng Ge, Weixin Mao +1
Currently, detecting 3D objects in Bird's-Eye-View (BEV) is superior to other 3D detectors for autonomous driving and robotics. However, transforming image features into BEV necess…
ThunderNet: Towards Real-time Generic Object Detection
Zheng Qin, Zeming Li, Zhaoning Zhang +4
Real-time generic object detection on mobile platforms is a crucial but challenging computer vision task. However, previous CNN-based detectors suffer from enormous computational c…
OTA: Optimal Transport Assignment for Object Detection
Zheng Ge, Songtao Liu, Zeming Li +2
Recent advances in label assignment in object detection mainly seek to independently define positive/negative training samples for each ground-truth (gt) object. In this paper, we…
An Experimental Study on Exploring Strong Lightweight Vision Transformers via Masked Image Modeling Pre-Training
Jin Gao, Shubo Lin, Shaoru Wang +7
Masked image modeling (MIM) pre-training for large-scale vision transformers (ViTs) has enabled promising downstream performance on top of the learned self-supervised ViT features.…
DBQ-SSD: Dynamic Ball Query for Efficient 3D Object Detection
Jinrong Yang, Lin Song, Songtao Liu +6
Many point-based 3D detectors adopt point-feature sampling strategies to drop some points for efficient inference. These strategies are typically based on fixed and handcrafted rul…
A Closer Look at Self-Supervised Lightweight Vision Transformers
Shaoru Wang, Jin Gao, Zeming Li +2
Self-supervised learning on large-scale Vision Transformers (ViTs) as pre-training methods has achieved promising downstream performance. Yet, how much these pre-training paradigms…
Distribution Alignment: A Unified Framework for Long-tail Visual Recognition
Songyang Zhang, Zeming Li, Shipeng Yan +2
Despite the recent success of deep neural networks, it remains challenging to effectively model the long-tail class distribution in visual recognition tasks. To address this proble…
Self-EMD: Self-Supervised Object Detection without ImageNet
Songtao Liu, Zeming Li, Jian Sun
In this paper, we propose a novel self-supervised representation learning method, Self-EMD, for object detection. Our method directly trained on unlabeled non-iconic image dataset…
DSPNet: Towards Slimmable Pretrained Networks based on Discriminative Self-supervised Learning
Shaoru Wang, Zeming Li, Jin Gao +2
Self-supervised learning (SSL) has achieved promising downstream performance. However, when facing various resource budgets in real-world applications, it costs a huge computation…
BorderDet: Border Feature for Dense Object Detection
Han Qiu, Yuchen Ma, Zeming Li +2
Dense object detectors rely on the sliding-window paradigm that predicts the object over a regular grid of image. Meanwhile, the feature maps on the point of the grid are adopted t…
Dynamic Scale Training for Object Detection
Yukang Chen, Peizhen Zhang, Zeming Li +5
We propose a Dynamic Scale Training paradigm (abbreviated as DST) to mitigate scale variation challenge in object detection. Previous strategies like image pyramid, multi-scale tra…
ESPO: Early-Stopping Proximal Policy Optimization
Zihang Li, Rui Zhou, Yingcheng Shi +8
When a large language model under reinforcement learning commits a wrong reasoning step early in a trajectory, standard algorithms force it to keep generating until the maximum hor…
DINO-Tok: Adapting DINO for Visual Tokenizers
Mingkai Jia, Mingxiao Li, Zhijian Shu +12
Recent advances in visual generation have emphasized the importance of Latent Generative Models (LGMs), which critically depend on effective visual tokenizers to bridge pixels and…
Multi-modal Relation Distillation for Unified 3D Representation Learning
Huiqun Wang, Yiping Bao, Panwang Pan +4
Recent advancements in multi-modal pre-training for 3D point clouds have demonstrated promising results by aligning heterogeneous features across 3D shapes and their corresponding…
BEVDepth: Acquisition of Reliable Depth for Multi-view 3D Object Detection
Yinhao Li, Zheng Ge, Guanyi Yu +5
In this research, we propose a new 3D object detector with a trustworthy depth estimation, dubbed BEVDepth, for camera-based Bird's-Eye-View (BEV) 3D object detection. Our work is…
MatrixVT: Efficient Multi-Camera to BEV Transformation for 3D Perception
Hongyu Zhou, Zheng Ge, Zeming Li +1
This paper proposes an efficient multi-camera to Bird's-Eye-View (BEV) view transformation method for 3D perception, dubbed MatrixVT. Existing view transformers either suffer from…
Fully Convolutional Networks for Panoptic Segmentation
Yanwei Li, Hengshuang Zhao, Xiaojuan Qi +4
In this paper, we present a conceptually simple, strong, and efficient framework for panoptic segmentation, called Panoptic FCN. Our approach aims to represent and predict foregrou…
AutoAssign: Differentiable Label Assignment for Dense Object Detection
Benjin Zhu, Jianfeng Wang, Zhengkai Jiang +4
Determining positive/negative samples for object detection is known as label assignment. Here we present an anchor-free detector named AutoAssign. It requires little human knowledg…
HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations
Peng Dai, Yang Zhang, Tao Liu +5
It is especially challenging to achieve real-time human motion tracking on a standalone VR Head-Mounted Display (HMD) such as Meta Quest and PICO. In this paper, we propose HMD-Pos…
Light-Head R-CNN: In Defense of Two-Stage Object Detector
Zeming Li, Chao Peng, Gang Yu +3
In this paper, we first investigate why typical two-stage methods are not as fast as single-stage, fast detectors like YOLO and SSD. We find that Faster R-CNN and R-FCN perform an…
Generalized Few-Shot Object Detection without Forgetting
Zhibo Fan, Yuchen Ma, Zeming Li +1
Recently few-shot object detection is widely adopted to deal with data-limited situations. While most previous works merely focus on the performance on few-shot categories, we clai…
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
Taewhan Kim, Hojin Bae, Zeming Li +4
Visual actionable affordance has emerged as a transformative approach in robotics, focusing on perceiving interaction areas prior to manipulation. Traditional methods rely on pixel…
StreamYOLO: Real-time Object Detection for Streaming Perception
Jinrong Yang, Songtao Liu, Zeming Li +2
The perceptive models of autonomous driving require fast inference within a low latency for safety. While existing works ignore the inevitable environmental changes after processin…
STS: Surround-view Temporal Stereo for Multi-view 3D Detection
Zengran Wang, Chen Min, Zheng Ge +4
Learning accurate depth is essential to multi-view 3D object detection. Recent approaches mainly learn depth from monocular images, which confront inherent difficulties due to the…
YOLOX: Exceeding YOLO Series in 2021
Zheng Ge, Songtao Liu, Feng Wang +2
In this report, we present some experienced improvements to YOLO series, forming a new high-performance detector -- YOLOX. We switch the YOLO detector to an anchor-free manner and…
MetaAnchor: Learning to Detect Objects with Customized Anchors
Tong Yang, Xiangyu Zhang, Zeming Li +2
We propose a novel and flexible anchor mechanism named MetaAnchor for object detection frameworks. Unlike many previous detectors model anchors via a predefined manner, in MetaAnch…
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
Chenguo Lin, Panwang Pan, Bangbang Yang +2
Recent advancements in 3D content generation from text or a single image struggle with limited high-quality 3D datasets and inconsistency from 2D multi-view generation. We introduc…
MegDet: A Large Mini-Batch Object Detector
Chao Peng, Tete Xiao, Zeming Li +5
The improvements in recent CNN-based object detection works, from R-CNN [11], Fast/Faster R-CNN [10, 31] to recent Mask R-CNN [14] and RetinaNet [24], mainly come from new network,…
Quality Matters: Embracing Quality Clues for Robust 3D Multi-Object Tracking
Jinrong Yang, En Yu, Zeming Li +2
3D Multi-Object Tracking (MOT) has achieved tremendous achievement thanks to the rapid development of 3D object detection and 2D MOT. Recent advanced works generally employ a serie…
Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection
Benjin Zhu, Zhengkai Jiang, Xiangxin Zhou +2
This report presents our method which wins the nuScenes3D Detection Challenge [17] held in Workshop on Autonomous Driving(WAD, CVPR 2019). Generally, we utilize sparse 3D convoluti…
EqCo: Equivalent Rules for Self-supervised Contrastive Learning
Benjin Zhu, Junqiang Huang, Zeming Li +2
In this paper, we propose EqCo (Equivalent Rules for Contrastive Learning) to make self-supervised learning irrelevant to the number of negative samples in the contrastive learning…
Workshop on Autonomous Driving at CVPR 2021: Technical Report for Streaming Perception Challenge
Songyang Zhang, Lin Song, Songtao Liu +4
In this report, we introduce our real-time 2D object detection system for the realistic autonomous driving scenario. Our detector is built on a newly designed YOLO model, called YO…
Momentum^2 Teacher: Momentum Teacher with Momentum Statistics for Self-Supervised Learning
Zeming Li, Songtao Liu, Jian Sun
In this paper, we present a novel approach, Momentum Teacher, for student-teacher based self-supervised learning. The approach performs momentum update on both network weights…
Learning Dynamic Routing for Semantic Segmentation
Yanwei Li, Lin Song, Yukang Chen +4
Recently, numerous handcrafted and searched networks have been applied for semantic segmentation. However, previous works intend to handle inputs with various scales in pre-defined…
Dense Teacher: Dense Pseudo-Labels for Semi-supervised Object Detection
Hongyu Zhou, Zheng Ge, Songtao Liu +4
To date, the most powerful semi-supervised object detectors (SS-OD) are based on pseudo-boxes, which need a sequence of post-processing with fine-tuned hyper-parameters. In this wo…
GMM: Delving into Gradient Aware and Model Perceive Depth Mining for Monocular 3D Detection
Weixin Mao, Jinrong Yang, Zheng Ge +5
Depth perception is a crucial component of monoc-ular 3D detection tasks that typically involve ill-posed problems. In light of the success of sample mining techniques in 2D object…
Dexbotic: Open-Source Vision-Language-Action Toolbox
Bin Xie, Erjin Zhou, Fan Jia +36
In this paper, we present Dexbotic, an open-source Vision-Language-Action (VLA) model toolbox based on PyTorch. It aims to provide a one-stop VLA research service for professionals…
NoiseAR: AutoRegressing Initial Noise Prior for Diffusion Models
Zeming Li, Xiangyue Liu, Xiangyu Zhang +2
Diffusion models have emerged as powerful generative frameworks, creating data samples by progressively denoising an initial random state. Traditionally, this initial state is samp…
Rebalanced Siamese Contrastive Mining for Long-Tailed Recognition
Zhisheng Zhong, Jiequan Cui, Zeming Li +3
Deep neural networks perform poorly on heavily class-imbalanced datasets. Given the promising performance of contrastive learning, we propose Rebalanced Siamese Contrastive Mining…
Fully Convolutional Networks for Panoptic Segmentation with Point-based Supervision
Yanwei Li, Hengshuang Zhao, Xiaojuan Qi +6
In this paper, we present a conceptually simple, strong, and efficient framework for fully- and weakly-supervised panoptic segmentation, called Panoptic FCN. Our approach aims to r…
Joint COCO and Mapillary Workshop at ICCV 2019: COCO Instance Segmentation Challenge Track
Zeming Li, Yuchen Ma, Yukang Chen +2
In this report, we present our object detection/instance segmentation system, MegDetV2, which works in a two-pass fashion, first to detect instances then to obtain segmentation. Ou…
End-to-End Object Detection with Fully Convolutional Network
Jianfeng Wang, Lin Song, Zeming Li +3
Mainstream object detectors based on the fully convolutional network has achieved impressive performance. While most of them still need a hand-designed non-maximum suppression (NMS…
Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets
Weiyu Li, Xuanyang Zhang, Zheng Sun +15
While generative artificial intelligence has advanced significantly across text, image, audio, and video domains, 3D generation remains comparatively underdeveloped due to fundamen…
Fine-Grained Dynamic Head for Object Detection
Lin Song, Yanwei Li, Zhengkai Jiang +4
The Feature Pyramid Network (FPN) presents a remarkable approach to alleviate the scale variance in object representation by performing instance-level assignments. Nevertheless, th…