3 papers
cs.CL2026
Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect
Ely Hahami, Ishaan Sinha, Lavik Jain
Can small language models detect and report on perturbations their own internal activations? We investigate this question through the lens of activation steering: injecting concept…
cs.AI2026
Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs
Ely Hahami, Ishaan Sinha, Lavik Jain +2
Can large language models introspect, that is, accurately detect perturbations to their own internal states? We systematically investigate this question using activation steering i…
cs.AI2025
NiceWebRL: a Python library for human subject experiments with reinforcement learning environments
Wilka Carvalho, Vikram Goddla, Ishaan Sinha +2
We present NiceWebRL, a research tool that enables researchers to use machine reinforcement learning (RL) environments for online human subject experiments. NiceWebRL is a Python l…