5 papers
Local Causal Attribution of Chain-of-Thought Reasoning
Dennis Wei, Yannis Belkhiter, Erik Miehling +1
Understanding the causal structure of a language model's thought process is a problem of significant importance for both transparency and safety. In this work, we take a local appr…
Multi-component Causal Tracing in Large Language Models
Zirui Yan, Dennis Wei, Dmitriy A. Katz +2
Causal tracing systematically intervenes on a large language model's (LLM's) internal representations to uncover and quantify the causal pathways linking specific inputs or computa…
Entropy-Aware On-Policy Distillation of Language Models
Woogyeol Jin, Taywon Min, Yongjin Yang +5
On-policy distillation is a promising approach for transferring knowledge between language models, where a student learns from dense token-level signals along its own trajectories.…
AI Steerability 360: A Toolkit for Steering Large Language Models
Erik Miehling, Karthikeyan Natesan Ramamurthy, Praveen Venkateswaran +10
The AI Steerability 360 toolkit is an extensible, open-source Python library for steering LLMs. Steering abstractions are designed around four model control surfaces: input (modifi…
Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
Hyo Jin Do, Rachel Ostrand, Werner Geyer +3
Large language models (LLMs) are susceptible to generating inaccurate or false information, often referred to as "hallucinations" or "confabulations." While several technical advan…