10 papers
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR
Yongshi Ye, Liang Zhang, Yidong Chen +2
Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR r…
PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation
Yongshi Ye, Biao Fu, Chongxuan Huang +2
Multi-domain machine translation (MDMT) requires more than fluent generation: it demands domain-sensitive translation decisions such as domain disambiguation, terminology control,…
Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation
Yongshi Ye, Biao Fu, Chongxuan Huang +2
Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by human translators' ability to adapt reas…
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
Xinjie Chen, Biao Fu, Jing Wu +4
Reinforcement learning with verifiable rewards (RLVR) enhances the reasoning of large language models (LLMs), but standard RLVR often depends on human-annotated answers or carefull…
UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory
Yongshi Ye, Hui Jiang, Feihu Jiang +7
Self-evolving memory serves as the trainable parameters for Large Language Models (LLMs)-based agents, where extraction (distilling insights from experience) and management (updati…
From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment
Chongxuan Huang, Yongshi Ye, Biao Fu +2
Large language models (LLMs) have demonstrated remarkable multilingual capabilities, however, how to evaluate cross-lingual alignment remains underexplored. Existing alignment benc…