collaborators

11 papers

cs.AI2026

FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents

Dongyi Lv, Fushun E, Aichen Cai +8

Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering tasks, retaining those that pass the tests, a…

cs.AI2026

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

Zhixiang Liang, Yifei Liu, Yidan Huang +5

Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long…

cs.AI2026

SearchMaster: Grounded and Regulated Self-Play for Search Agents

Wentao Tan, Qiong Cao, Jiaqi Wang +1

Training LLM-based search agents requires high-quality search data: tasks that demand genuine multi-hop retrieval and trajectories that use search tools effectively. Existing pipel…

cs.SD2026

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

Yinhao Bai, Jinming Chen, Yafeng Chen +26

We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligenc…

cs.LG2026

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning

Wenpu Liu, Yuqi Xu, Weichu Xie +8

Reinforcement Learning from Verifiable Rewards (RLVR) typically samples multiple responses per prompt and assigns binary rewards based on individual correctness, yet the collective…

cs.LG2026

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning

Ziyue Wang, Aomufei Yuan, Yongfu Zhu +10

Reinforcement Learning from Verifiable Rewards (RLVR) has become the dominant approach for improving mathematical reasoning in large language models, yet current methods reduce eac…