activity
20172026
most citedEfficientFormer: Vision Transformers at MobileNet Speed

257 citations · 585 across the 28 of their papers we have counts for

collaborators
Showing 2020Show all

6 papers · 1 filter

cs.CV2020★ 5 cited

Achieving Real-Time LiDAR 3D Object Detection on a Mobile Device

Pu Zhao, Wei Niu, Geng Yuan +7

3D object detection is an important task, especially in the autonomous driving application domain. However, it is challenging to support the real-time performance with the limited…

cs.LG2020

NPAS: A Compiler-aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile Acceleration

Zhengang Li, Geng Yuan, Wei Niu +13

With the increasing demand to efficiently deploy DNNs on mobile edge devices, it becomes much more important to reduce unnecessary computation and increase the execution speed. Pri…

cs.CV2020

ClickTrain: Efficient and Accurate End-to-End Deep Learning Training via Fine-Grained Architecture-Preserving Pruning

Chengming Zhang, Geng Yuan, Wei Niu +8

Convolutional neural networks (CNNs) are becoming increasingly deeper, wider, and non-linear because of the growing demand on prediction accuracy and analysis quality. The wide and…

cs.CL2020

Real-Time Execution of Large-scale Language Models on Mobile

Wei Niu, Zhenglun Kong, Geng Yuan +7

Pre-trained large-scale language models have increasingly demonstrated high accuracy on many natural language processing (NLP) tasks. However, the limited weight storage and comput…

cs.CV2020

YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-Design

Yuxuan Cai, Hongjia Li, Geng Yuan +5

The rapid development and wide utilization of object detection techniques have aroused attention on both accuracy and speed of object detectors. However, the current state-of-the-a…

cs.LG2020★ 14 cited

SS-Auto: A Single-Shot, Automatic Structured Weight Pruning Framework of DNNs with Ultra-High Efficiency

Zhengang Li, Yifan Gong, Xiaolong Ma +6

Structured weight pruning is a representative model compression technique of DNNs for hardware efficiency and inference accelerations. Previous works in this area leave great space…