2 papers
cs.LG2025
Preemptive Detection and Steering of LLM Misalignment via Latent Reachability
Sathwik Karnik, Somil Bansal
Large language models (LLMs) are now ubiquitous in everyday tools, raising urgent safety concerns about their tendency to generate harmful content. The dominant safety approach --…
cs.RO2024
Embodied Red Teaming for Auditing Robotic Foundation Models
Sathwik Karnik, Zhang-Wei Hong, Nishant Abhangi +5
Language-conditioned robot models have the potential to enable robots to perform a wide range of tasks based on natural language instructions. However, assessing their safety and e…