Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
DeepInsight II: One Trace from Benchmark to Robot
Siyi Li, Yuchen Kang, Wuliang Wang +4
Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized harnesses, while the embodied layers on whic…
cs.AI2026
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack
Siyi Li, Chunyu Sun, Jiahao Zhang +6
Evaluating a Physical AI stack spans operators that differ by more than three orders of magnitude -- from a single foundation-model decoding step to thousands of physics ticks of w…