3 papers
cs.AI2025
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
Cameron Tice, Philipp Alexander Kreer, Nathan Helm-Burger +7
Capability evaluations play a crucial role in assessing and regulating frontier AI systems. The effectiveness of these evaluations faces a significant challenge: strategic underper…
cs.CL2025
Noise Injection Systemically Degrades Large Language Model Safety Guardrails
Prithviraj Singh Shahani, Kaveh Eskandari Miandoab, Matthias Scheutz
Safety guardrails in large language models (LLMs) are a critical component in preventing harmful outputs. Yet, their resilience under perturbation remains poorly understood. In thi…
cs.RO2025
Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture
Hong Lu, Hengxu Li, Prithviraj Singh Shahani +2
Vision-language-action (VLA) models hold promise as generalist robotics solutions by translating visual and linguistic inputs into robot actions, yet they lack reliability due to t…