3 papers
cs.CV2026
Cumulative Consensus Score: Label-Free and Model-Agnostic Evaluation of Object Detectors in Deployment
Avinaash Manoharan, Xiangyu Yin, Domenik Helm +1
Evaluating object detection models in deployment is challenging because ground-truth annotations are rarely available. We introduce the Cumulative Consensus Score (CCS), a label-fr…
cs.RO2026
Validating Generalist Robots with Situation Calculus and STL Falsification
Changwen Li, Rongjie Yan, Chih-Hong Cheng +1
Generalist robots are becoming a reality, capable of interpreting natural language instructions and executing diverse operations. However, their validation remains challenging beca…
cs.SE2025
Rethinking Code Review Workflows with LLM Assistance: An Empirical Study
Fannar Steinn Aðalsteinsson, Björn Borgar Magnússon, Mislav Milicevic +2
Code reviews are a critical yet time-consuming aspect of modern software development, increasingly challenged by growing system complexity and the demand for faster delivery. This…