5 papers · 1 filter
Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation
Xuan-Phi Nguyen, Shrey Pandit, Zeyu Leo Liu +5
On-policy self-distillation aims to improve upon reinforcement learning from verifiable rewards (RLVR) by providing token-level scores derived from privileged information, such as…
Procedural Memory Distillation: Online Reflection for Self-Improving Language Models
Ye Liu, Srijan Bansal, Bo Pang +6
Reinforcement learning with verifiable rewards (RLVR), along with recent selfdistillation variants such as SDPO, evaluates each rollout against a verifier and updates the policy fr…
Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning
Xuanzhi Feng, Zhengyang Li, Zeyu Liu +6
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced Large Language Model (LLM) reasoning; however, it faces a fundamental optimization instability: uni…
LingLanMiDian: Systematic Evaluation of LLMs on TCM Knowledge and Clinical Reasoning
Rui Hua, Yu Wei, Zixin Shu +33
Large language models (LLMs) are advancing rapidly in medical NLP, yet Traditional Chinese Medicine (TCM) with its distinctive ontology, terminology, and reasoning patterns require…
WebGuard: Building a Generalizable Guardrail for Web Agents
Boyuan Zheng, Zeyi Liao, Scott Salisbury +8
The rapid development of autonomous web agents powered by Large Language Models (LLMs), while greatly elevating efficiency, exposes the frontier risk of taking unintended or harmfu…