collaborators

5 papers

cs.CL2026

Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution

Xiaoou Liu, Tiejin Chen, Dengjia Zhang +3

Large Language Models have achieved strong performance on reasoning tasks with objective answers by generating step-by-step solutions, but diagnosing where a multi-step reasoning t…

cs.RO2026

ContextFlow: Hierarchical Task-State Alignment for Long-Horizon Embodied Agents

Shuhan Guo, Kun Zhang, Haifei Liu +4

Long-horizon embodied agents increasingly delegate navigation, search, approach, and manipulation to specialist executors. As these executors become stronger, the main bottleneck s…

cs.LG2026

SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards

Dengjia Zhang, Xiaoou Liu, Lu Cheng +3

Large language models (LLMs) are increasingly deployed as multi-step decision-making agents, where effective reward design is essential for guiding learning. Although recent work e…

cs.LG2025

MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios

Xuantang Xiong, Ni Mu, Runpeng Xie +8

Model-based reinforcement learning (MBRL) is a crucial approach to enhance the generalization capabilities and improve the sample efficiency of RL algorithms. However, current MBRL…

cs.LG2025

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks

Pengbo Shen, Yaqing Wang, Ni Mu +8

Evaluating large language models (LLMs) in complex decision-making is essential for advancing AI's ability for strategic planning and real-time adaptation. However, existing benchm…