collaborators

11 papers

cs.AI2026

AndroidReality: How Far Are Mobile Agents from the Real World?

Xiaoou Liu, Longchao Da, Hanyang Chen +2

Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-world deployment due to environm…

cs.CL2026

Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution

Xiaoou Liu, Tiejin Chen, Dengjia Zhang +3

Large Language Models have achieved strong performance on reasoning tasks with objective answers by generating step-by-step solutions, but diagnosing where a multi-step reasoning t…

cs.AI2026

The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective

Xiaoou Liu, Tiejin Chen, Weibo Li +2

Foundation model agents are increasingly deployed for real-world decision-making, but suffer from the sim-to-real gap. While robotics and classical control have mature frameworks t…

cs.CL2026

Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering

Tiejin Chen, Longchao Da, Xiaoou Liu +1

Uncertainty Quantification (UQ) is widely regarded as the primary safeguard for deploying Large Language Models (LLMs) in high-stakes domains. However, we argue that the field suff…

cs.CL2026

LangMARL: Natural Language Multi-Agent Reinforcement Learning

Huaiyuan Yao, Longchao Da, Xiaoou Liu +3

Large language model (LLM) agents struggle to autonomously evolve coordination strategies in dynamic environments, largely because coarse global outcomes obscure the causal signals…

cs.LG2026

Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment

Tiejin Chen, Xiaoou Liu, Vishnu Nandam +2

Preference-based alignment like Reinforcement Learning from Human Feedback (RLHF) learns from pairwise preferences, yet the labels are often noisy and inconsistent. Existing uncert…