6 papers
Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
Fangxu Yu, Tao Feng, Dehai Min +6
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs…
Weak-to-Strong On-Policy Distillation
Fangxu Yu, Weijia Xu, Michael Xu +2
On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabiliti…
Rushes: A Human Preference Dataset for Pluralistic Alignment
Michael Xu, Jorge Leandro, Sudha Rao +5
We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collected through a game interface…
GFlowRL: Scaling Distribution-Matching RL to Large Language Models
Xiaodong Liu, Michael Xu, Jack W. Stokes +3
Generative Flow Networks (GFlowNets) offer a promising alternative to reward-maximizing reinforcement learning (RL) for large reasoning models, encouraging diverse reasoning paths…
MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents
Tamil Sudaravan Mohan Doss, Michael Xu, Sudha Rao +2
We present MineNPC-Task, a user-authored benchmark and evaluation harness for testing memory-aware, mixed-initiative LLM agents in open-world Minecraft. Rather than relying on synt…
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
Yihao Guo, Haocheng Bian, Liutong Zhou +16
With the deployment of Large Language Models (LLMs) in interactive applications, online malicious intent detection has become increasingly critical. However, existing approaches fa…