119 citations · 595 across the 28 of their papers we have counts for
8 papers · 1 filter
COMPOSER: Compositional Reasoning of Group Activity in Videos with Keypoint-Only Modality
Honglu Zhou, Asim Kadav, Aviv Shamsian +6
Group Activity Recognition detects the activity collectively performed by a group of actors, which requires compositional reasoning of actors and objects. We approach the task by m…
A Simple Long-Tailed Recognition Baseline via Vision-Language Model
Teli Ma, Shijie Geng, Mengmeng Wang +5
The visual world naturally exhibits a long-tailed distribution of open classes, which poses great challenges to modern visual systems. Existing approaches either perform class re-b…
Audio-Visual Scene-Aware Dialog and Reasoning using Audio-Visual Transformers with Joint Student-Teacher Learning
Ankit P. Shah, Shijie Geng, Peng Gao +5
In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at bo…
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
Peng Gao, Shijie Geng, Renrui Zhang +5
Large-scale contrastive vision-language pre-training has shown significant progress in visual representation learning. Unlike traditional visual systems trained by a fixed set of d…
Dense Contrastive Visual-Linguistic Pretraining
Lei Shi, Kai Shuang, Shijie Geng +5
Inspired by the success of BERT, several multimodal representation learning approaches have been proposed that jointly represent image and text. These approaches achieve superior p…
Counterfactual Evaluation for Explainable AI
Yingqiang Ge, Shuchang Liu, Zelong Li +6
While recent years have witnessed the emergence of various explainable methods in machine learning, to what degree the explanations really represent the reasoning process behind th…