collaborators

5 papers

cs.LG2026

Local Causal Attribution of Chain-of-Thought Reasoning

Dennis Wei, Yannis Belkhiter, Erik Miehling +1

Understanding the causal structure of a language model's thought process is a problem of significant importance for both transparency and safety. In this work, we take a local appr…

cs.LG2026

Multi-component Causal Tracing in Large Language Models

Zirui Yan, Dennis Wei, Dmitriy A. Katz +2

Causal tracing systematically intervenes on a large language model's (LLM's) internal representations to uncover and quantify the causal pathways linking specific inputs or computa…

cs.LG2026

Entropy-Aware On-Policy Distillation of Language Models

Woogyeol Jin, Taywon Min, Yongjin Yang +5

On-policy distillation is a promising approach for transferring knowledge between language models, where a student learns from dense token-level signals along its own trajectories.…

cs.CL2026

AI Steerability 360: A Toolkit for Steering Large Language Models

Erik Miehling, Karthikeyan Natesan Ramamurthy, Praveen Venkateswaran +10

The AI Steerability 360 toolkit is an extensible, open-source Python library for steering LLMs. Steering abstractions are designed around four model control surfaces: input (modifi…

cs.HC2025

Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators

Hyo Jin Do, Rachel Ostrand, Werner Geyer +3

Large language models (LLMs) are susceptible to generating inaccurate or false information, often referred to as "hallucinations" or "confabulations." While several technical advan…