collaborators

19 papers

cs.AI2026

AndroidReality: How Far Are Mobile Agents from the Real World?

Xiaoou Liu, Longchao Da, Hanyang Chen +2

Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-world deployment due to environm…

cs.LG2026

DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

Hanyang Chen, Anirudh Satheesh, Longchao Da +1

Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains. In this paper, we consid…

cs.CL2026

Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution

Xiaoou Liu, Tiejin Chen, Dengjia Zhang +3

Large Language Models have achieved strong performance on reasoning tasks with objective answers by generating step-by-step solutions, but diagnosing where a multi-step reasoning t…

cs.CV2026

ShadeBench: A Benchmark Dataset for Building Shade Simulation in Sustainable Society

Longchao Da, Mithun Shivakoti, Xiangrui Liu +3

Urban heat exposure is becoming an increasingly critical challenge due to the intensifying urban heat island effect. Fine-grained shade patterns, especially those induced by urban…

cs.CL2026

Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering

Tiejin Chen, Longchao Da, Xiaoou Liu +1

Uncertainty Quantification (UQ) is widely regarded as the primary safeguard for deploying Large Language Models (LLMs) in high-stakes domains. However, we argue that the field suff…

cs.CL2026

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

Haolin Chen, Deon Metelski, Leon Qi +30

End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large l…