104 citations · 498 across the 22 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023★ 12 cited
Aligning Large Multimodal Models with Factually Augmented RLHF
Zhiqing Sun, Sheng Shen, Shengcao Cao +9
Large Multimodal Models (LMM) are built across modalities and the misalignment between two modalities can result in "hallucination", generating textual outputs that are not grounde…
cs.CV2020
Rethinking Transformer-based Set Prediction for Object Detection
Zhiqing Sun, Shengcao Cao, Yiming Yang +1
DETR is a recently proposed Transformer-based method which views object detection as a set prediction problem and achieves state-of-the-art performance but demands extra-long train…
cs.CV2020
VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
Jingzhou Liu, Wenhu Chen, Yu Cheng +4
We introduce a new task, Video-and-Language Inference, for joint multimodal understanding of video and text. Given a video clip with aligned subtitles as premise, paired with a nat…