5 citations · 6 across the 5 of their papers we have counts for
7 papers
Difficulty-Estimated Policy Optimization
Yu Zhao, Fan Jiang, Tianle Liu +4
Recent advancements in Large Reasoning Models (LRMs), exemplified by DeepSeek-R1, have underscored the potential of scaling inference-time compute through Group Relative Policy Opt…
A State-Transition Framework for Efficient LLM Reasoning
Liang Zhang, Yu Zhao, Longyue Wang +4
While Long Chain-of-Thought (CoT) reasoning significantly improves Large Language Models (LLMs) performance on complex reasoning tasks, the substantial computational and memory cos…
Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models
Bo Zeng, Chenyang Lyu, Sinuo Liu +14
Instruction-following capability has become a major ability to be evaluated for Large Language Models (LLMs). However, existing datasets, such as IFEval, are either predominantly m…
The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks
Minghao Wu, Weixuan Wang, Sinuo Liu +7
As large language models (LLMs) continue to advance in linguistic capabilities, robust multilingual evaluation has become essential for promoting equitable technological progress.…
Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models
Huifeng Yin, Yu Zhao, Minghao Wu +9
Large Reasoning Models(LRMs) such as OpenAI o1 and DeepSeek-R1 have shown remarkable reasoning capabilities by scaling test-time compute and generating long Chain-of-Thought(CoT).…
Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement
Lingfeng Ming, Bo Zeng, Chenyang Lyu +17
Large Language Models (LLMs) have achieved remarkable progress in recent years; however, their excellent performance is still largely limited to major world languages, primarily En…