3 papers
cs.MM2026
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
Yang-Hao Zhou, Haitian Li, Rexar Lin +12
Recent advances in text-to-audio-video (T2AV) generation have enabled models to synthesize audio-visual videos with multi-participant dialogues. However, existing evaluation benchm…
cs.LG2025
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
Yuyang Xu, Yi Cheng, Haochao Ying +5
Test-time scaling has proven effective in further enhancing the performance of pretrained Large Language Models (LLMs). However, mainstream post-training methods (i.e., reinforceme…
cs.CL2025
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
Renjun Hu, Yi Cheng, Libin Meng +4
The rapid advancement of large language models (LLMs) has opened new possibilities for their adoption as evaluative judges. This paper introduces Themis, a fine-tuned LLM judge tha…