collaborators

6 papers

cs.LG2026

A Unifying Lens on Reward Uncertainty in RLHF

Ely Hahami, Yoel Zimmermann, Ray Zhou +1

Reinforcement learning from human feedback (RLHF) is bottlenecked by reward hacking, where the policy exploits errors in a proxy reward model (RM) and produces high RM scores witho…

cs.CL2026

Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs

Xu Pan, Ely Hahami, Jingxuan Fan +2

Large language models (LLMs) are often used in environments where facts evolve, yet factual knowledge updates via fine-tuning on unstructured text often suffer from 1) reliance on…

cs.CL2026

Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect

Ely Hahami, Ishaan Sinha, Lavik Jain

Can small language models detect and report on perturbations their own internal activations? We investigate this question through the lens of activation steering: injecting concept…

cs.CL2026

User-Assistant Bias in LLMs

Xu Pan, Jingxuan Fan, Zidi Xiong +3

Modern large language models (LLMs) are typically trained and deployed using structured role tags (e.g. system, user, assistant, tool) that explicitly mark the source of each piece…

cs.AI2026

Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs

Ely Hahami, Ishaan Sinha, Lavik Jain +2

Can large language models introspect, that is, accurately detect perturbations to their own internal states? We systematically investigate this question using activation steering i…

cs.CL2025

Memorization and Knowledge Injection in Gated LLMs

Xu Pan, Ely Hahami, Zechen Zhang +1

Large Language Models (LLMs) currently struggle to sequentially add new memories and integrate new knowledge. These limitations contrast with the human ability to continuously lear…