4 papers
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process
Zican Hu, Xuyang Hu, Yiming Liu +10
Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi-turn generation via reinforcement learni…
VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes
Jingru Chen, Yiming Liu, Mingtao Chen +5
Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such scores do not necessarily imp…
EternalMath: A Living Benchmark of Frontier Mathematics that Evolves with Human Discovery
Jicheng Ma, Guohua Wang, Xinhua Feng +3
Current evaluations of mathematical reasoning in large language models (LLMs) are dominated by static benchmarks, either derived from competition-style problems or curated through…
Q-Save: Towards Scoring and Attribution for Generated Video Evaluation
Xiele Wu, Zicheng Zhang, Mingtao Chen +7
Evaluating AI-generated video (AIGV) quality hinges on three crucial dimensions: visual quality, dynamic quality, and text-video alignment. While numerous evaluation datasets and a…