collaborators

7 papers

cs.AI2026

CaveAgent: Transforming LLMs into Stateful Runtime Operators

Maohao Ran, Zhenglin Wan, Cooper Lin +21

LLM-based agents are increasingly capable of complex task execution, yet current agentic systems remain constrained by text-centric paradigms that struggle with long-horizon tasks…

cs.AI2026

Don't Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents

Chubin Zhang, Zhenglin Wan, Xingrui Yu +5

Tool-augmented agents are typically evaluated by their gains under reliable external feedback. Yet these gains leave open a key counterfactual: when feedback is unreliable, would t…

cs.AI2026

Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention

Chubin Zhang, Zhenglin Wan, Xingrui Yu +5

Runtime oversight for LLM agents is commonly framed as scalar risk prediction: estimate failure likelihood, confidence, or uncertainty, then intervene once the score crosses a thre…

cs.LG2026

Adversarial Dual On-Policy Distillation from Expressive Teacher

Zhenglin Wan, Jingxuan Wu, Xingrui Yu +5

Learning from demonstrations in embodied control is often cast as behavioral cloning, and recent diffusion or flow-matching policies improve this paradigm by modeling multi-modal e…

cs.LG2026

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning

Zhenglin Wan, Jingxuan Wu, Xingrui Yu +4

Flow Matching (FM) has shown remarkable ability in modeling complex distributions and achieves strong performance in offline imitation learning for cloning expert behaviors. Howeve…

cs.AI2026

Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching

Jingxuan Wu, Zhenglin Wan, Xingrui Yu +4

Flow-based text-to-image models follow deterministic trajectories, making it costly to explore diverse modes under limited sampling budgets. Existing approaches to improving divers…