works on

From the 1 of 7 linked papers with an AI index.

activity
20242026
collaborators

7 papers

cs.LG2026

Cluster-Weighted EDMD

Lorenzo Tomaz, Judd Rosenblatt, Flavio Kicis +2

The paper proposes Cluster-Weighted EDMD, a method that learns both a soft partition of the state space and separate EDMD operators for each cluster, improving prediction accuracy…

cs.LG2026

Modular Pretraining Enables Access Control

Ethan Roland, Murat Cubuktepe, Erick Martinez +8

AI developers face a dual-use dilemma. An AI capability that helps one user cure a disease can help another synthesize one. This dilemma could be resolved with access control, limi…

cs.LG2026

Endogenous Resistance to Activation Steering in Language Models

Alex McKenzie, Keenan Pepper, Stijn Servaes +6

Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.g., ``wait, that's not right'') and continuing on-t…

cs.CL2026

Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs

Keenan Pepper, Alex McKenzie, Florin Pop +6

Self-interpretation methods prompt language models to describe their own internal states, but remain unreliable due to hyperparameter sensitivity. We show that training lightweight…

cs.CL2025

Large Language Models Report Subjective Experience Under Self-Referential Processing

Cameron Berg, Diogo de Lucena, Judd Rosenblatt

Large language models sometimes produce structured, first-person descriptions that explicitly reference awareness or subjective experience. To better understand this behavior, we i…

cs.CL2025

Momentum Point-Perplexity Mechanics in Large Language Models

Lorenzo Tomaz, Judd Rosenblatt, Thomas Berry Jones +1

We take a physics-based approach to studying how the internal hidden states of large language models change from token to token during inference. Across 20 open-source transformer…