6 papers
Don't Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents
Chubin Zhang, Zhenglin Wan, Xingrui Yu +5
Tool-augmented agents are typically evaluated by their gains under reliable external feedback. Yet these gains leave open a key counterfactual: when feedback is unreliable, would t…
Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention
Chubin Zhang, Zhenglin Wan, Xingrui Yu +5
Runtime oversight for LLM agents is commonly framed as scalar risk prediction: estimate failure likelihood, confidence, or uncertainty, then intervene once the score crosses a thre…
Training Diffusion Policies via Prior-Mapping Co-Evolution
Chubin Zhang, Zhenglin Wan, Feng Chen +7
Reinforcement learning (RL) faces a persistent tension: policies that are stable to optimize (e.g., Gaussians) are often too simple to represent the multimodal action distributions…
Adversarial Dual On-Policy Distillation from Expressive Teacher
Zhenglin Wan, Jingxuan Wu, Xingrui Yu +5
Learning from demonstrations in embodied control is often cast as behavioral cloning, and recent diffusion or flow-matching policies improve this paradigm by modeling multi-modal e…
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
Zhenglin Wan, Jingxuan Wu, Xingrui Yu +4
Flow Matching (FM) has shown remarkable ability in modeling complex distributions and achieves strong performance in offline imitation learning for cloning expert behaviors. Howeve…
LexChain: Modeling Legal Reasoning Chains for Chinese Tort Case Analysis
Huiyuan Xie, Chenyang Li, Huining Zhu +4
Legal reasoning is a fundamental component of legal analysis and decision-making. Existing computational approaches to legal reasoning predominantly rely on generic reasoning frame…