2 papers
cs.CL2026
LLM Layers Immediately Correct Each Other
Arjun Patrawala, Jiahai Feng, Erik Jones +1
Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linear, semantically meaningful feat…
cs.CL2025
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
Erik Jones, Arjun Patrawala, Jacob Steinhardt
Humans often rely on subjective natural language to direct language models (LLMs); for example, users might instruct the LLM to write an enthusiastic blogpost, while developers mig…