143 citations · 606 across the 35 of their papers we have counts for
42 papers · 1 filter
Aligning Large Multimodal Models with Factually Augmented RLHF
Zhiqing Sun, Sheng Shen, Shengcao Cao +9
Large Multimodal Models (LMM) are built across modalities and the misalignment between two modalities can result in "hallucination", generating textual outputs that are not grounde…
Learning Situation Hyper-Graphs for Video Question Answering
Aisha Urooj Khan, Hilde Kuehne, Bo Wu +5
Answering questions about complex situations in videos requires not only capturing the presence of actors, objects, and their relations but also the evolution of these relationship…
AutoGPart: Intermediate Supervision Search for Generalizable 3D Part Segmentation
Xueyi Liu, Xiaomeng Xu, Anyi Rao +2
Training a generalizable 3D part segmentation network is quite challenging but of great importance in real-world applications. To tackle this problem, some works design task-specif…
When Does Contrastive Learning Preserve Adversarial Robustness from Pretraining to Finetuning?
Lijie Fan, Sijia Liu, Pin-Yu Chen +2
Contrastive learning (CL) can learn generalizable feature representations and achieve the state-of-the-art performance of downstream tasks by finetuning a linear classifier on top…
TSM: Temporal Shift Module for Efficient and Scalable Video Understanding on Edge Device
Ji Lin, Chuang Gan, Kuan Wang +1
The explosive growth in video streaming requires video understanding at high accuracy and low computation cost. Conventional 2D CNNs are computationally cheap but cannot capture te…
Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning
Shaobo Min, Qi Dai, Hongtao Xie +3
Cross-modal correlation provides an inherent supervision for video unsupervised representation learning. Existing methods focus on distinguishing different video clips by visual an…