activity
20222024
most citedAIM: Adapting Image Models for Efficient Video Action Recognition

62 citations · 79 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AI2024

Law of the Weakest Link: Cross Capabilities of Large Language Models

Ming Zhong, Aston Zhang, Xuewei Wang +14

The development and evaluation of Large Language Models (LLMs) have largely focused on individual capabilities. However, this overlooks the intersection of multiple abilities acros…

cs.CL20231 cited

Vcc: Scaling Transformers to 128K Tokens or More by Prioritizing Important Tokens

Zhanpeng Zeng, Cole Hawkins, Mingyi Hong +4

Transformers are central in modern natural language processing and computer vision applications. Despite recent works devoted to reducing the quadratic cost of such models (as a fu…

cs.CV202362 cited

AIM: Adapting Image Models for Efficient Video Action Recognition

Taojiannan Yang, Yi Zhu, Yusheng Xie +3

Recent vision transformer based video models mostly follow the ``image pre-training then finetuning" paradigm and have achieved great success on multiple video benchmarks. However,…

cs.CV202210 cited

Partial and Asymmetric Contrastive Learning for Out-of-Distribution Detection in Long-Tailed Recognition

Haotao Wang, Aston Zhang, Yi Zhu +4

Existing out-of-distribution (OOD) detection methods are typically benchmarked on training sets with balanced class distributions. However, in real-world applications, it is common…

cs.LG20226 cited

Removing Batch Normalization Boosts Adversarial Training

Haotao Wang, Aston Zhang, Shuai Zheng +3

Adversarial training (AT) defends deep neural networks against adversarial attacks. One challenge that limits its practical application is the performance degradation on clean samp…