5 papers
Efficient and Robust Video Defense Framework against 3D-field Personalized Talking Face
Rui-qing Sun, Xingshan Yao, Tian Lan +6
State-of-the-art 3D-field video-referenced Talking Face Generation (TFG) methods synthesize high-fidelity personalized talking-face videos in real time by modeling 3D geometry and…
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
Tian Lan, Yang-Hao Zhou, Zi-Ao Ma +8
Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these gener…
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
Zi-Ao Ma, Tian Lan, Rong-Cheng Tu +5
The rapid progress in diffusion-based text-to-image (T2I) generation has created an urgent need for interpretable automatic evaluation methods that can assess the quality of genera…
Subtopic-aware View Sampling and Temporal Aggregation for Long-form Document Matching
Youchao Zhou, Heyan Huang, Zhijing Wu +2
Long-form document matching aims to judge the relevance between two documents and has been applied to various scenarios. Most existing works utilize hierarchical or long context mo…
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
Zi-Ao Ma, Tian Lan, Rong-Cheng Tu +6
We present a systematic investigation of Multi-modal Retrieval Augmented Multi-modal Generation (MRAG), a novel task that enables foundation models to process multi-modal web c…