4 papers
Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering
Hao Wang, Jiuzhou Lei, Dayou Li +5
Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise undesired at deployment. For example, a polic…
Learning Actionable Manipulation Recovery via Counterfactual Failure Synthesis
Dayou Li, Jiuzhou Lei, Hao Wang +6
While recent foundation models have significantly advanced robotic manipulation, these systems still struggle to autonomously recover from execution errors. Current failure-learnin…
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
Hongze Tan, Zihan Wang, Jianfei Pan +7
Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarse-grained cr…
DEEPMED: Building a Medical DeepResearch Agent via Multi-hop Med-Search Data and Turn-Controlled Agentic Training & Inference
Zihan Wang, Hao Wang, Shi Feng +6
Medical reasoning models remain constrained by parametric knowledge and are thus susceptible to forgetting and hallucinations. DeepResearch (DR) models ground outputs in verifiable…