Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
FailBench: How Reliable are VLMs at Judging Robot Task Success?
Zaruhi Navasardyan, Tatul Danielyan, Hrant Davtyan
Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but existing benchmarks offer limited evidence of cross-domain generalization. We intro…
cs.RO2026
A Statistical Audit of Physical AI Benchmark Redundancy
Zaruhi Navasardyan, Hrant Davtyan
Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and the relationship between benchmarks unme…