8 papers
HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains
Zheng Li, Mao Zheng, Mingyang Song +1
General-purpose machine translation benchmarks such as FLORES-200 have reached a saturation regime on Chinese-English pairs, where modern large language models cluster within a nar…
IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following
Mingrui Sun, Mao Zheng, Zheng Li +1
Modern translation workflows demand more than semantic equivalence. Users routinely require models to preserve JSON or HTML schemas, honor curated glossaries, disambiguate with pro…
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
Wenjie Yang, Mao Zheng, Mingyang Song +2
Large language models (LLMs) have recently demonstrated remarkable capabilities in machine translation (MT). However, most advanced MT-specific LLMs heavily rely on external superv…
HY-MT1.5 Technical Report
Mao Zheng, Zheng Li, Tao Chen +2
In this report, we introduce our latest translation models, HY-MT1.5-1.8B and HY-MT1.5-7B, a new family of machine translation models developed through a holistic training framewor…
FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient Training R1-like Reasoning Models
Mingyang Song, Mao Zheng, Zheng Li +4
Improving training efficiency continues to be one of the primary challenges in large-scale Reinforcement Learning (RL). In this paper, we investigate how context length and the com…
Hunyuan-MT Technical Report
Mao Zheng, Zheng Li, Bingxin Qu +4
In this report, we introduce Hunyuan-MT-7B, our first open-source multilingual translation model, which supports bidirectional translation across 33 major languages and places a sp…