19 papers
AndroidReality: How Far Are Mobile Agents from the Real World?
Xiaoou Liu, Longchao Da, Hanyang Chen +2
Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-world deployment due to environm…
DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning
Hanyang Chen, Anirudh Satheesh, Longchao Da +1
Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains. In this paper, we consid…
ShadeBench: A Benchmark Dataset for Building Shade Simulation in Sustainable Society
Longchao Da, Mithun Shivakoti, Xiangrui Liu +3
Urban heat exposure is becoming an increasingly critical challenge due to the intensifying urban heat island effect. Fine-grained shade patterns, especially those induced by urban…
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
Tiejin Chen, Longchao Da, Xiaoou Liu +1
Uncertainty Quantification (UQ) is widely regarded as the primary safeguard for deploying Large Language Models (LLMs) in high-stakes domains. However, we argue that the field suff…
LangMARL: Natural Language Multi-Agent Reinforcement Learning
Huaiyuan Yao, Longchao Da, Xiaoou Liu +3
Large language model (LLM) agents struggle to autonomously evolve coordination strategies in dynamic environments, largely because coarse global outcomes obscure the causal signals…
Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous Driving
Longchao Da, David Isele, Hua Wei +1
Being able to anticipate the motion of surrounding agents is essential for the safe operation of autonomous driving systems in dynamic situations. While various methods have been p…