3 papers
cs.AI2025
FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect Generalization
Shibo Hong, Jiahao Ying, Haiyuan Liang +4
Evaluating open-ended outputs of Multimodal Large Language Models has become a bottleneck as model capabilities, task diversity, and modality rapidly expand. Existing ``MLLM-as-a-J…
cs.CL2025
Less Data Less Tokens: Multilingual Unification Learning for Efficient Test-Time Reasoning in LLMs
Kang Chen, Mengdi Zhang, Yixin Cao
This paper explores the challenges of test-time scaling of large language models (LLMs), regarding both the data and inference efficiency. We highlight the diversity of multi-lingu…
cs.LG2025
BNPO: Beta Normalization Policy Optimization
Changyi Xiao, Mengdi Zhang, Yixin Cao
Recent studies, including DeepSeek-R1 and Kimi-k1.5, have demonstrated that reinforcement learning with rule-based, binary-valued reward functions can significantly enhance the rea…