papers

Publications (10)

cs.CV2020

AFDet: Anchor Free One Stage 3D Object Detection

Runzhou Ge, Zhuangzhuang Ding, Yihan Hu +4

High-efficiency point cloud 3D object detection operated on embedded systems is important for many robotics applications including autonomous driving. Most previous works try to so…

cs.CV2020

1st Place Solutions for Waymo Open Dataset Challenges -- 2D and 3D Tracking

Yu Wang, Sijia Chen, Li Huang +4

This technical report presents the online and real-time 2D and 3D multi-object tracking (MOT) algorithms that reached the 1st places on both Waymo Open Dataset 2D tracking and 3D t…

cs.CV2020

1st Place Solution for Waymo Open Dataset Challenge -- 3D Detection and Domain Adaptation

Zhuangzhuang Ding, Yihan Hu, Runzhou Ge +4

In this technical report, we introduce our winning solution "HorizonLiDAR3D" for the 3D detection track and the domain adaptation track in Waymo Open Dataset Challenge at CVPR 2020…

cs.CV2021

Real-Time Anchor-Free Single-Stage 3D Detection with IoU-Awareness

Runzhou Ge, Zhuangzhuang Ding, Yihan Hu +4

In this report, we introduce our winning solution to the Real-time 3D Detection and also the "Most Efficient Model" in the Waymo Open Dataset Challenges at CVPR 2021. Extended from…

cs.CV2018

MAC: Mining Activity Concepts for Language-based Temporal Localization

Runzhou Ge, Jiyang Gao, Kan Chen +1

We address the problem of language-based temporal localization in untrimmed videos. Compared to temporal localization with fixed categories, this problem is more challenging as the…

cs.CV2024

WOMD-LiDAR: Raw Sensor Dataset Benchmark for Motion Forecasting

Kan Chen, Runzhou Ge, Hang Qiu +12

Widely adopted motion forecasting datasets substitute the observed sensory inputs with higher-level abstractions such as 3D boxes and polylines. These sparse shapes are inferred th…

cs.CV2018

Motion-Appearance Co-Memory Networks for Video Question Answering

Jiyang Gao, Runzhou Ge, Kan Chen +1

Video Question Answering (QA) is an important task in understanding video temporal structure. We observe that there are three unique attributes of video QA compared with image QA:…

cs.CV2022

AFDetV2: Rethinking the Necessity of the Second Stage for Object Detection from Point Clouds

Yihan Hu, Zhuangzhuang Ding, Runzhou Ge +4

There have been two streams in the 3D detection from point clouds: single-stage methods and two-stage methods. While the former is more computationally efficient, the latter usuall…

cs.CV2024

MoST: Multi-modality Scene Tokenization for Motion Prediction

Norman Mu, Jingwei Ji, Zhenpei Yang +11

Many existing motion prediction approaches rely on symbolic perception outputs to generate agent trajectories, such as bounding boxes, road graph information and traffic lights. Th…

cs.CV2020

2nd Place Solution for Waymo Open Dataset Challenge -- 2D Object Detection

Sijia Chen, Yu Wang, Li Huang +4

A practical autonomous driving system urges the need to reliably and accurately detect vehicles and persons. In this report, we introduce a state-of-the-art 2D object detection sys…