3 papers
cs.AI2026
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR
Yongshi Ye, Liang Zhang, Yidong Chen +2
Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR r…
cs.CL2026
From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation
Yu Pan, Yang Hou, Xiongfei Wu +4
Compositional speech-to-speech translation (S2ST) systems built upon speech large language models (SpeechLLMs) have recently shown promising performance. However, existing S2ST sys…
cs.CL2025
LLMs Can Achieve High-quality Simultaneous Machine Translation as Efficiently as Offline
Biao Fu, Minpeng Liao, Kai Fan +4
When the complete source sentence is provided, Large Language Models (LLMs) perform excellently in offline machine translation even with a simple prompt "Translate the following se…