activity
20162023
most citedThe Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision

143 citations · 606 across the 35 of their papers we have counts for

collaborators
Showing cs.CVShow all

42 papers · 1 filter

cs.CV2023★ 12 cited

Aligning Large Multimodal Models with Factually Augmented RLHF

Zhiqing Sun, Sheng Shen, Shengcao Cao +9

Large Multimodal Models (LMM) are built across modalities and the misalignment between two modalities can result in "hallucination", generating textual outputs that are not grounde…

cs.CV2023★ 4 cited

Learning Situation Hyper-Graphs for Video Question Answering

Aisha Urooj Khan, Hilde Kuehne, Bo Wu +5

Answering questions about complex situations in videos requires not only capturing the presence of actors, objects, and their relations but also the evolution of these relationship…

cs.CV2022

AutoGPart: Intermediate Supervision Search for Generalizable 3D Part Segmentation

Xueyi Liu, Xiaomeng Xu, Anyi Rao +2

Training a generalizable 3D part segmentation network is quite challenging but of great importance in real-world applications. To tackle this problem, some works design task-specif…

cs.CV2021★ 6 cited

When Does Contrastive Learning Preserve Adversarial Robustness from Pretraining to Finetuning?

Lijie Fan, Sijia Liu, Pin-Yu Chen +2

Contrastive learning (CL) can learn generalizable feature representations and achieve the state-of-the-art performance of downstream tasks by finetuning a linear classifier on top…

cs.CV2021★ 1 cited

TSM: Temporal Shift Module for Efficient and Scalable Video Understanding on Edge Device

Ji Lin, Chuang Gan, Kuan Wang +1

The explosive growth in video streaming requires video understanding at high accuracy and low computation cost. Conventional 2D CNNs are computationally cheap but cannot capture te…

cs.CV2021★ 5 cited

Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning

Shaobo Min, Qi Dai, Hongtao Xie +3

Cross-modal correlation provides an inherent supervision for video unsupervised representation learning. Existing methods focus on distinguishing different video clips by visual an…