Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
Xinrun Xu, Pi Bu, Ye Wang +7
Although Vision Language Models (VLMs) exhibit strong perceptual abilities and impressive visual reasoning, they struggle with attention to detail and precise action planning in co…
cs.AI2024
Cradle: Empowering Foundation Agents Towards General Computer Control
Weihao Tan, Wentao Zhang, Xinrun Xu +25
Despite the success in specific scenarios, existing foundation agents still struggle to generalize across various virtual scenarios, mainly due to the dramatically different encaps…