6 papers
A Unifying Lens on Reward Uncertainty in RLHF
Ely Hahami, Yoel Zimmermann, Ray Zhou +1
Reinforcement learning from human feedback (RLHF) is bottlenecked by reward hacking, where the policy exploits errors in a proxy reward model (RM) and produces high RM scores witho…
Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs
Xu Pan, Ely Hahami, Jingxuan Fan +2
Large language models (LLMs) are often used in environments where facts evolve, yet factual knowledge updates via fine-tuning on unstructured text often suffer from 1) reliance on…
Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect
Ely Hahami, Ishaan Sinha, Lavik Jain
Can small language models detect and report on perturbations their own internal activations? We investigate this question through the lens of activation steering: injecting concept…
User-Assistant Bias in LLMs
Xu Pan, Jingxuan Fan, Zidi Xiong +3
Modern large language models (LLMs) are typically trained and deployed using structured role tags (e.g. system, user, assistant, tool) that explicitly mark the source of each piece…
Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs
Ely Hahami, Ishaan Sinha, Lavik Jain +2
Can large language models introspect, that is, accurately detect perturbations to their own internal states? We systematically investigate this question using activation steering i…
Memorization and Knowledge Injection in Gated LLMs
Xu Pan, Ely Hahami, Zechen Zhang +1
Large Language Models (LLMs) currently struggle to sequentially add new memories and integrate new knowledge. These limitations contrast with the human ability to continuously lear…