4 papers · 1 filter
Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate
Chenxi Liu, Yanshuo Chen, Ruibo Chen +3
The reasoning abilities of large language models (LLMs) have been substantially improved by reinforcement learning with verifiable rewards (RLVR). At test time, collaborative reaso…
LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling
Tong Zheng, Haolin Liu, Chengsong Huang +10
Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during inference. However, existing TTS…
Explore Data Left Behind in Reinforcement Learning for Reasoning Language Models
Chenxi Liu, Junjie Liang, Yuqi Jia +4
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an effective approach for improving the reasoning abilities of large language models (LLMs). The Group Relative…
A Watermark for Order-Agnostic Language Models
Ruibo Chen, Yihan Wu, Yanshuo Chen +3
Statistical watermarking techniques are well-established for sequentially decoded language models (LMs). However, these techniques cannot be directly applied to order-agnostic LMs,…