2 papers
cs.AI2026
DeepInsight II: One Trace from Benchmark to Robot
Siyi Li, Yuchen Kang, Wuliang Wang +4
Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized harnesses, while the embodied layers on whic…
cs.RO2025
Domain-Conditioned Scene Graphs for State-Grounded Task Planning
Jonas Herzog, Jiangpin Liu, Yue Wang
Recent robotic task planning frameworks have integrated large multimodal models (LMMs) such as GPT-4o. To address grounding issues of such models, it has been suggested to split th…