8 papers
GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents
Xiao Zhou, Siyue Zhang, Yilun Zhao +4
Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces, but existing evaluations confound interaction modality with d…
Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation
Siyi Gu, Jialin Chen, Sophia Zhou +2
Post-training of reasoning language models is commonly driven by supervised distillation and reinforcement learning with verifiable rewards. Distillation often relies on chain-of-t…
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
Tianyu Liu, Allen Xin Wang, Antonia Panescu +30
AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchma…
Quantifying Faithful Confidence Expression in Large Reasoning Models
Areeb Gani, Asal Meskin, Gabrielle Kaili-May Liu +1
Reliable uncertainty communication is critical to the trustworthiness of LLMs, yet faithful calibration (FC)--the alignment between models' intrinsic and (linguistically) expressed…
Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence?
Gabrielle Kaili-May Liu, Arman Cohan
LLMs' linguistically expressed confidence should faithfully reflect their intrinsic uncertainty. While recent work shows LLMs struggle to use epistemic markers (e.g., "it is likely…
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
Jinbiao Wei, Qianran Ma, Yilun Zhao +4
We present OpenComputer, a verifier-grounded framework for constructing verifiable software worlds for computer-use agents. OpenComputer integrates four components: (1) app-specifi…