6 papers
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
Zihan Dong, Rui Qian, Qishi Zhan +3
Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We identify the main cause as expe…
How Benchmarks Mis-Score Computer-Use Agents
Zihan Dong, Zhiyuan Ma, Zekun Wang +5
Computer-use agents (CUA) are being deployed to browse the web and operate desktop software, yet their benchmark scores are still commonly produced by brittle scripted oracles. A s…
Discontinuous Prior-Mode Sections and the Geometry of Ambiguity in Intrinsic Image Decomposition
Ziheng Chen, Liangchen Liu, Qishi Zhan +2
The viral 2015 photograph known as "The Dress" divides observers into two camps because it is ambiguous: the same image colors can be explained either as a blue-black surface under…
Stress Amplified Resilience: ESG and Joint Fragility in Equity Markets
Minxuan Hu, Jiayu Yi, Ziheng Chen +2
Market stress rarely harms investors through one channel alone. Losses, volatility spikes, and deteriorating tradability often arrive together. We examine whether ESG is associated…
A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning
Qishi Zhan, Minxuan Hu, Liang He +2
In limited-data settings, a single endpoint mean of an evaluation metric such as the Continuous Ranked Probability Score (CRPS) is itself a random variable, yet it is routinely rep…
Unstable Rankings in Bayesian Deep Learning Evaluation
Qishi Zhan, Minxuan Hu, Guansu Wang +2
Standard evaluations of Bayesian deep learning methods assume that metric estimates are reliable, but we show this assumption fails under data scarcity. Method rankings are not onl…