1.8k citations · 3.6k across the 86 of their papers we have counts for
20 papers · 1 filter
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
Yuanhao Xiong, Long Zhao, Boqing Gong +5
Existing video-language pre-training methods primarily focus on instance-level alignment between video clips and captions via global contrastive learning but neglect rich fine-grai…
Scaling Up Dataset Distillation to ImageNet-1K with Constant Memory
Justin Cui, Ruochen Wang, Si Si +1
Dataset Distillation is a newly emerging area that aims to distill large datasets into much smaller and highly informative synthetic ones to accelerate training and reduce storage.…
Generalizing Few-Shot NAS with Gradient Matching
Shoukang Hu, Ruochen Wang, Lanqing Hong +3
Efficient performance estimation of architectures drawn from large search spaces is essential to Neural Architecture Search. One-Shot methods tackle this challenge by training one…
Temporal Shuffling for Defending Deep Action Recognition Models against Adversarial Attacks
Jaehui Hwang, Huan Zhang, Jun-Ho Choi +2
Recently, video-based action recognition methods using convolutional neural networks (CNNs) achieve remarkable recognition performance. However, there is still lack of understandin…
Can Vision Transformers Perform Convolution?
Shanda Li, Xiangning Chen, Di He +1
Several recent studies have demonstrated that attention-based networks, such as Vision Transformer (ViT), can outperform Convolutional Neural Networks (CNNs) on several computer vi…
Adversarial Attack across Datasets
Yunxiao Qin, Yuanhao Xiong, Jinfeng Yi +2
Existing transfer attack methods commonly assume that the attacker knows the training set (e.g., the label set, the input size) of the black-box victim models, which is usually unr…