2 papers
cs.MM2026
RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning
Ruoxuan Zhang, Qiyun Zheng, Siyu Wu +13
Vision-Language Models (VLMs) are widely used for visual understanding, yet current evaluation protocols fail to assess whether these capabilities are grounded in physical reasonin…
cs.RO2025
Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Intervention
Yanbo Mao, Jianlong Fu, Ruoxuan Zhang +2
Vision-Language-Action (VLA) models have enabled notable progress in general-purpose robotic manipulation, yet their learned policies often exhibit variable execution quality. We a…