9 papers
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
Haitian Li, Yanghao Zhou, Heyan Huang +15
In recent years, Multi-Talker Audio-Video Generation (MTAVG) models have shown promising performance on fundamental metrics such as lip-sync and audio-visual alignment. However, th…
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
Yang-Hao Zhou, Haitian Li, Rexar Lin +12
Recent advances in text-to-audio-video (T2AV) generation have enabled models to synthesize audio-visual videos with multi-participant dialogues. However, existing evaluation benchm…
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Surveys
Guo-Biao Zhang, Ding-Yuan Liu, Da-Yi Wu +5
The rapid development of automated survey generation technology has made it increasingly important to establish a comprehensive benchmark to evaluate the quality of generated surve…
MMWOZ: Building Multimodal Agent for Task-oriented Dialogue
Pu-Hai Yang, Heyan Huang, Heng-Da Xu +3
Task-oriented dialogue systems have garnered significant attention due to their conversational ability to accomplish goals, such as booking airline tickets for users. Traditionally…
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
Tian Lan, Yang-Hao Zhou, Zi-Ao Ma +8
Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these gener…
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
Zi-Ao Ma, Tian Lan, Rong-Cheng Tu +5
The rapid progress in diffusion-based text-to-image (T2I) generation has created an urgent need for interpretable automatic evaluation methods that can assess the quality of genera…