9 papers
Trace-Based On-Policy Distillation for Masked Diffusion Language Models
Haolin Ren, Ziyang Huang, Chenhao Yuan +2
Diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. However, reasoning-oriented post-training for dLLMs remains challenging. Supervise…
A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions
Jianghan Shen, Siqi Luo, Yue Li +9
Policy gradient algorithms for language models optimize the same objective , which has exactly two factors: the trajectory probability…
MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline
Jiyao Liu, Jianghan Shen, Sida Song +19
Recent advances in deep research systems enable large language models to retrieve, synthesize, and reason over large-scale external knowledge. In medicine, developing clinical guid…
R3A: Reinforced Reasoning for Relevance Assessment for RAG in User-Generated Content Platforms
Xiaowei Yuan, Lei Jin, Haoxin Zhang +6
Retrieval-augmented generation (RAG) plays a critical role in user-generated content (UGC) platforms, but its effectiveness critically depends on accurate query-document relevance…
WideSeek: Advancing Wide Research via Multi-Agent Scaling
Ziyang Huang, Haolin Ren, Xiaowei Yuan +6
Search intelligence is evolving from Deep Research to Wide Research, a paradigm essential for retrieving and synthesizing comprehensive information under complex constraints in par…
Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns
Xiang Li, Haiyang Yu, Xinghua Zhang +6
Process Reward Models (PRMs) are crucial in complex reasoning and problem-solving tasks (e.g., LLM agents with long-horizon decision-making) by verifying the correctness of each in…