5 papers
Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement
Xin Wang, Liangtai Sun, Yaoming Zhu +8
Existing code-generation benchmarks score a single mapping from a complete prompt to a one-shot output. However, real web development is different. Users seldom write a full spec a…
WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
Kaining Ying, Hengrui Hu, Siyu Ren +6
Interactive world models are advancing rapidly, yet existing benchmarks cover only part of the required competencies, leaving no unified standard for systematic evaluation. To fill…
Q-REAL: Towards Realism and Plausibility Evaluation for AI-Generated Content
Shushi Wang, Zicheng Zhang, Chunyi Li +7
Quality assessment of AI-generated content is crucial for evaluating model capability and guiding model optimization. However, most existing quality assessment datasets and models…
UniHetero: Could Generation Enhance Understanding for Vision-Language-Model at Large Data Scale?
Fengjiao Chen, Minhao Jing, Weitao Lu +3
Vision-language large models are moving toward the unification of visual understanding and visual generation tasks. However, whether generation can enhance understanding is still u…
EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans
Yingjie Zhou, Xilei Zhu, Siyu Ren +11
Speech-driven Talking Human (TH) generation, commonly known as "Talker," currently faces limitations in multi-subject driving capabilities. Extending this paradigm to "Multi-Talker…