activity
20172026
most citedRecDCL: Dual Contrastive Learning for Recommendation

61 citations · 159 across the 33 of their papers we have counts for

collaborators
Showing 2024Show all

12 papers · 1 filter

cs.CV2024★ 1 cited

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models

Tao Wu, Yong Zhang, Xiaodong Cun +6

Zero-shot customized video generation has gained significant attention due to its substantial application potential. Existing methods rely on additional models to extract and injec…

cs.CV2024

DOGR: Towards Versatile Visual Document Grounding and Referring

Yinan Zhou, Yuxin Chen, Haokun Lin +6

With recent advances in Multimodal Large Language Models (MLLMs), grounding and referring capabilities have gained increasing attention for achieving detailed understanding and fle…

cs.AI2024

mRAG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Tao Zhang, Ziqi Zhang, Zongyang Ma +10

Advanced Multimodal Large Language Models (MLLMs) struggle with recent Knowledge-based Visual Question Answering (VQA) tasks, such as INFOSEEK and Encyclopedic-VQA, due to their li…

cs.CV2024★ 2 cited

Taming Rectified Flow for Inversion and Editing

Jiangshan Wang, Junfu Pu, Zhongang Qi +6

Rectified-flow-based diffusion transformers like FLUX and OpenSora have demonstrated outstanding performance in the field of image and video generation. Despite their robust genera…

cs.CV2024

E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Ye Liu, Zongyang Ma, Zhongang Qi +3

Recent advances in Video Large Language Models (Video-LLMs) have demonstrated their great potential in general-purpose video understanding. To verify the significance of these mode…

cs.CV2024

SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses

Chaolei Tan, Zihang Lin, Junfu Pu +7

Video grounding is a fundamental problem in multimodal content understanding, aiming to localize specific natural language queries in an untrimmed video. However, current video gro…