7 papers
Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models
Bo Zeng, Chenyang Lyu, Sinuo Liu +14
Instruction-following capability has become a major ability to be evaluated for Large Language Models (LLMs). However, existing datasets, such as IFEval, are either predominantly m…
Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation
Xintong Wang, Jingheng Pan, Yixiao Liu +8
Vision-Language Translation (VLT) is a challenging task that requires accurately recognizing multilingual text embedded in images and translating it into the target language with t…
Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models
Huifeng Yin, Yu Zhao, Minghao Wu +9
Large Reasoning Models(LRMs) such as OpenAI o1 and DeepSeek-R1 have shown remarkable reasoning capabilities by scaling test-time compute and generating long Chain-of-Thought(CoT).…
TransBench: Benchmarking Machine Translation for Industrial-Scale Applications
Haijun Li, Tianqi Shi, Zifu Shang +13
Machine translation (MT) has become indispensable for cross-border communication in globalized industries like e-commerce, finance, and legal services, with recent advancements in…
The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks
Minghao Wu, Weixuan Wang, Sinuo Liu +7
As large language models (LLMs) continue to advance in linguistic capabilities, robust multilingual evaluation has become essential for promoting equitable technological progress.…
New Trends for Modern Machine Translation with Large Reasoning Models
Sinuo Liu, Chenyang Lyu, Minghao Wu +4
Recent advances in Large Reasoning Models (LRMs), particularly those leveraging Chain-of-Thought reasoning (CoT), have opened brand new possibility for Machine Translation (MT). Th…