activity
20182024
most citedLLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

119 citations · 595 across the 28 of their papers we have counts for

collaborators
Showing 2021Show all

8 papers · 1 filter

cs.CV2021

COMPOSER: Compositional Reasoning of Group Activity in Videos with Keypoint-Only Modality

Honglu Zhou, Asim Kadav, Aviv Shamsian +6

Group Activity Recognition detects the activity collectively performed by a group of actors, which requires compositional reasoning of actors and objects. We approach the task by m…

cs.CV2021★ 10 cited

A Simple Long-Tailed Recognition Baseline via Vision-Language Model

Teli Ma, Shijie Geng, Mengmeng Wang +5

The visual world naturally exhibits a long-tailed distribution of open classes, which poses great challenges to modern visual systems. Existing approaches either perform class re-b…

cs.CL2021

Audio-Visual Scene-Aware Dialog and Reasoning using Audio-Visual Transformers with Joint Student-Teacher Learning

Ankit P. Shah, Shijie Geng, Peng Gao +5

In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at bo…

cs.CV2021★ 113 cited

CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Peng Gao, Shijie Geng, Renrui Zhang +5

Large-scale contrastive vision-language pre-training has shown significant progress in visual representation learning. Unlike traditional visual systems trained by a fixed set of d…

cs.CV2021

Dense Contrastive Visual-Linguistic Pretraining

Lei Shi, Kai Shuang, Shijie Geng +5

Inspired by the success of BERT, several multimodal representation learning approaches have been proposed that jointly represent image and text. These approaches achieve superior p…

cs.CL2021★ 6 cited

Counterfactual Evaluation for Explainable AI

Yingqiang Ge, Shuchang Liu, Zelong Li +6

While recent years have witnessed the emergence of various explainable methods in machine learning, to what degree the explanations really represent the reasoning process behind th…