2 papers
cs.CV2026
SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments
Suryanarayana Reddy Yarrabothula, Manisha Chawla, Kunal Sinha +5
Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under…
cs.AI2026
Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge
Dhaval Patel, Chathurangi Shyalika, Suryanarayana Reddy Yarrabothula +4
Competition retrospectives are useful when they explain what a leaderboard measured, how hidden evaluation changed conclusions, and which design patterns were rewarded. We revisit…