5 citations · 24 across the 10 of their papers we have counts for
5 papers
FocusFormer: Focusing on What We Need via Architecture Sampler
Jing Liu, Jianfei Cai, Bohan Zhuang
Vision Transformers (ViTs) have underpinned the recent breakthroughs in computer vision. However, designing the architectures of ViTs is laborious and heavily relies on expert know…
An Efficient Spatio-Temporal Pyramid Transformer for Action Detection
Yuetian Weng, Zizheng Pan, Mingfei Han +2
The task of action detection aims at deducing both the action category and localization of the start and end moment for each action instance in a long, untrimmed video. While visio…
Sequential Person Recognition in Photo Albums with a Recurrent Network
Yao Li, Guosheng Lin, Bohan Zhuang +3
Recognizing the identities of people in everyday photos is still a very challenging problem for machine vision, due to non-frontal faces, changes in clothing, location, lighting an…
Attend in groups: a weakly-supervised deep learning framework for learning from web data
Bohan Zhuang, Lingqiao Liu, Yao Li +2
Large-scale datasets have driven the rapid development of deep neural networks for visual recognition. However, annotating a massive dataset is expensive and time-consuming. Web im…
Visual Tracking via Shallow and Deep Collaborative Model
Bohan Zhuang, Lijun Wang, Huchuan Lu
In this paper, we propose a robust tracking method based on the collaboration of a generative model and a discriminative classifier, where features are learned by shallow and deep…