5 papers
Provably Extracting the Features from a General Superposition
Allen Liu
It is widely believed that complex machine learning models generally encode features through linear representations. This is the foundational hypothesis behind a vast body of work…
Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
Ishaq Aden-Ali, Noah Golowich, Allen Liu +3
Training modern large language models (LLMs) has become a veritable smorgasbord of algorithms and datasets designed to elicit particular behaviors, making it critical to develop te…
Provably Learning from Modern Language Models via Low Logit Rank
Noah Golowich, Allen Liu, Abhishek Shetty
While modern language models and their inner workings are incredibly complex, recent work (Golowich, Liu & Shetty; 2025) has proposed a simple and potentially tractable abstraction…
Sequences of Logits Reveal the Low Rank Structure of Language Models
Noah Golowich, Allen Liu, Abhishek Shetty
A major problem in the study of large language models is to understand their inherent low-dimensional structure. We introduce an approach to study the low-dimensional structure of…
Are LLMs complicated ethical dilemma analyzers?
Jiashen, Du, Jesse Yao +2
One open question in the study of Large Language Models (LLMs) is whether they can emulate human ethical reasoning and act as believable proxies for human judgment. To investigate…