1.8k citations · 2k across the 7 of their papers we have counts for
15 papers
EAPruning: Evolutionary Pruning for Vision Transformers and CNNs
Qingyuan Li, Bo Zhang, Xiangxiang Chu
Structured pruning greatly eases the deployment of large neural networks in resource-constrained environments. However, current methods either involve strong domain expertise, requ…
YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications
Chuyi Li, Lulu Li, Hongliang Jiang +15
For years, the YOLO series has been the de facto industry-level standard for efficient object detection. The YOLO community has prospered overwhelmingly to enrich its use in a mult…
Modeling Motion with Multi-Modal Features for Text-Based Video Segmentation
Wangbo Zhao, Kai Wang, Xiangxiang Chu +3
Text-based video segmentation aims to segment the target object in a video based on a describing sentence. Incorporating motion information from optical flow maps with appearance a…
CCTrans: Simplifying and Improving Crowd Counting with Transformer
Ye Tian, Xiangxiang Chu, Hongpeng Wang
Most recent methods used for crowd counting are based on the convolutional neural network (CNN), which has a strong ability to extract local features. But CNN inherently fails in m…
Twins: Revisiting the Design of Spatial Attention in Vision Transformers
Xiangxiang Chu, Zhi Tian, Yuqing Wang +5
Very recently, a variety of vision transformer architectures for dense prediction tasks have been proposed and they show that the design of spatial attention is critical to their s…
AutoKWS: Keyword Spotting with Differentiable Architecture Search
Bo Zhang, Wenfeng Li, Qingyuan Li +3
Smart audio devices are gated by an always-on lightweight keyword spotting program to reduce power consumption. It is however challenging to design models that have both high accur…