12 papers
Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies
Murali Indukuri, Mohammad Eskandari, Sree Nitya Kollu +2
Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Vision-Language Models (VLMs) ar…
Autonomous VR-Based Risk Detection for Situational Awareness in Dangerous Settings
Mohammad Eskandari, Murali Krishna Varma Indukuri, Stephanie M. Lukin +1
In high-risk environments such as disaster response, situational awareness depends not only on detecting hazards but also on communicating them clearly to human operators. Vision L…
Limited Linguistic Diversity in Embodied AI Datasets
Selma Wanna, Agnes Luhtaru, Jonathan Salfity +4
Language plays a critical role in Vision-Language-Action (VLA) models, yet the linguistic characteristics of the datasets used to train and evaluate these systems remain poorly doc…
Query Disambiguation via Answer-Free Context: Doubling Performance on Humanity's Last Exam
Michael Majurski, Cynthia Matuszek
How carefully and unambiguously a question is phrased has a profound impact on the quality of the response, for Language Models (LMs) as well as people. While model capabilities co…
Grounding Synthetic Data Evaluations of Language Models in Unsupervised Document Corpora
Michael Majurski, Cynthia Matuszek
Language Models (LMs) continue to advance, improving response quality and coherence. Given Internet-scale training datasets, LMs have likely encountered much of what users may ask…
Scaling Patterns in Adversarial Alignment: Evidence from Multi-LLM Jailbreak Experiments
Samuel Nathanson, Rebecca Williams, Cynthia Matuszek
Large language models (LLMs) increasingly operate in multi-agent and safety-critical settings, raising open questions about how their vulnerabilities scale when models interact adv…