3 citations · 4 across the 28 of their papers we have counts for
5 papers · 1 filter
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
Mingyang Song, Mao Zheng
As test-time scaling becomes a pivotal research frontier in Large Language Models (LLMs) development, contemporary and advanced post-training methodologies increasingly focus on ex…
TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment
Zheng Li, Mao Zheng, Mingyang Song +1
Recently, deep reasoning large language models(LLMs) like DeepSeek-R1 have made significant progress in tasks such as mathematics and coding. Inspired by this, several studies have…
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
Wenjie Yang, Mao Zheng, Mingyang Song +2
Large language models (LLMs) have recently demonstrated remarkable capabilities in machine translation (MT). However, most advanced MT-specific LLMs heavily rely on external superv…
FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient Training R1-like Reasoning Models
Mingyang Song, Mao Zheng, Zheng Li +4
Improving training efficiency continues to be one of the primary challenges in large-scale Reinforcement Learning (RL). In this paper, we investigate how context length and the com…
GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs
Mingyang Song, Mao Zheng, Xuan Luo
Pairwise LLM-as-a-judge evaluation asks the judge to identify the \emph{better} of two candidate answers. We study a one-line modification that asks for the \emph{worse} answer ins…