11 papers
Localized Adaptation Reveals Distinct Learning Signatures in Transformers
Rebecca Ramnauth, Brian Scassellati
Transformer adaptation is typically distributed across model depth, even when the intended change is narrow. We investigate how adaptation site shapes what a model learns, how well…
The Attentional White Bear Effect in Transformer Language Models
Rebecca Ramnauth, Brian Scassellati
Instruction-based suppression is widely used to prevent language models from generating prohibited content, yet it remains unclear whether suppression reduces internal representati…
Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains
Rebecca Ramnauth, Drazen Brscic, Brian Scassellati
Foundation models are increasingly deployed in socially sensitive domains such as education, mental health, and caregiving, where failures are often cumulative and context-dependen…
Towards Zero-Knowledge Task Planning via a Language-based Approach
Liam Merz Hoffmeister, Brian Scassellati, Daniel Rakita
In this work, we introduce and formalize the Zero-Knowledge Task Planning (ZKTP) problem, i.e., formulating a sequence of actions to achieve some goal without task-specific knowled…
Open-Ended Goal Inference through Actions and Language for Human-Robot Collaboration
Debasmita Ghose, Oz Gitelson, Marynel Vazquez +1
To collaborate with humans, robots must infer goals that are often ambiguous, difficult to articulate, or not drawn from a fixed set. Prior approaches restrict inference to a prede…
I've Changed My Mind: Robots Adapting to Changing Human Goals during Collaboration
Debasmita Ghose, Oz Gitelson, Ryan Jin +3
For effective human-robot collaboration, a robot must align its actions with human goals, even as they change mid-task. Prior approaches often assume fixed goals, reducing goal pre…