9 papers
Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems
Prajjwal Gupta, Prasang Gupta, Vishal Bhutani +4
As agentic LLM systems move from prototypes to deployment across increasingly diverse domains, evaluating them has become both more important and more difficult. The challenge is n…
StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement
Junwon Seo, Sushant Veer, Ran Tian +6
Video world models (WMs) have shown promise for policy evaluation and improvement by imagining realistic future observations conditioned on ego-robot actions. While WMs can model d…
Tractable Uncertainty-Aware Meta-Learning
Young-Jin Park, Cesar Almecija, Apoorva Sharma +1
Meta-learning is a popular approach for learning new tasks with limited data by leveraging the commonalities among different tasks. However, meta-learned models can perform poorly…
Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
Zhenghao "Mark" Peng, Wenhao Ding, Yurong You +11
Recent reasoning-augmented Vision-Language-Action (VLA) models have improved the interpretability of end-to-end autonomous driving by generating intermediate reasoning traces. Yet…
The Case for Negative Data: From Crash Reports to Counterfactuals for Reasonable Driving
Jay Patrikar, Apoorva Sharma, Sushant Veer +3
Learning-based autonomous driving systems are trained mostly on incident-free data, offering little guidance near safety-performance boundaries. Real crash reports contain precisel…
Sim2Val: Leveraging Correlation Across Test Platforms for Variance-Reduced Metric Estimation
Rachel Luo, Heng Yang, Michael Watson +4
Learning-based robotic systems demand rigorous validation to assure reliable performance, but extensive real-world testing is often prohibitively expensive, and if conducted may st…