Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
Luyu Yang, Yutong Dai, An Yan +3
The physical world is not merely visual; it is governed by rigorous structural and procedural constraints. Yet, the evaluation of vision-language models (VLMs) remains heavily skew…
cs.AI2025
SCUBA: Salesforce Computer Use Benchmark
Yutong Dai, Krithika Ramakrishnan, Jing Gu +8
We introduce SCUBA, a benchmark designed to evaluate computer-use agents on customer relationship management (CRM) workflows within the Salesforce platform. SCUBA contains 300 task…