activity
20242026
collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Beyond Linear Steering: Unified Multi-Attribute Control for Language Models

Narmeen Oozeer, Luke Marks, Shreyans Jain +2

Controlling multiple behavioral attributes in large language models (LLMs) at inference time is a challenging problem due to interference between attributes and the limitations of…

cs.LG2025

DiFR: Inference Verification Despite Nondeterminism

Adam Karvonen, Daniel Reuter, Roy Rinberg +3

As demand for LLM inference grows, it is becoming increasingly important that providers and their customers can verify that inference processes are performed correctly, without err…

cs.LG2025

TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research

Abir Harrasse, Philip Quirke, Clement Neo +3

Mechanistic interpretability research faces a gap between analyzing simple circuits in toy tasks and discovering features in large models. To bridge this gap, we propose text-to-SQ…

cs.LG2025

Output Supervision Can Obfuscate the Chain of Thought

Jacob Drori, Luke Marks, Bryce Woodworth +2

OpenAI (2025) showed that training against a chain of thought (CoT) monitor can cause obfuscated CoTs, which contain bad behavior the monitor cannot detect. They proposed to keep C…

cs.LG2025

Interpreting Learned Feedback Patterns in Large Language Models

Luke Marks, Amir Abdullah, Clement Neo +4

Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferen…

cs.LG2024

Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders

Luke Marks, Alasdair Paren, David Krueger +1

Sparse Autoencoders (SAEs) have shown promise in improving the interpretability of neural network activations, but can learn features that are not features of the input, limiting t…