6 papers
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop
Yikun Fu, Bowen Fu, Zhenyu Wu +10
Computer use agents (CUAs) have advanced rapidly in desktop automation, and a growing number of users deploy CUAs such as OpenClaw on Mac Mini for always-on automation. However, ex…
CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming
Ruixun Liu, Lingyu Zhang, Lanxuan Xue +3
Humans can effortlessly reason about scenes across different viewpoints, yet it remains unclear whether Vision-Language Models (VLMs) possess similar cross-view spatial abilities.…
Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle
Jiayu Wang, Weijiang Lv, Bowen Fu +8
As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and eve…
Assessing the Effects of Monetary Shocks on Macroeconomic Stars: A SMUC-IV Framework
Bowen Fu, Chenghan Hou, Jan Prüser
This paper proposes a structural multivariate unobserved components model with external instrument (SMUC-IV) to investigate the effects of monetary policy shocks on key U.S. macroe…
Ovis-Image Technical Report
Guo-Hua Wang, Liangfu Cao, Tianyu Cui +8
We introduce , a 7B text-to-image model specifically optimized for high-quality text rendering, designed to operate efficiently under stringent computational c…
ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks
Ruixun Liu, Bowen Fu, Jiayi Song +7
Ultra-high-resolution (UHR) remote sensing (RS) images offer rich fine-grained information but also present challenges in effective processing. Existing dynamic resolution and toke…