4 papers
Position: Token Taxes Can Mitigate AI's Economic Risks
Lucas Irwin, Tung-Yu Wu, Fazl Barez
AI-driven automation threatens to erode government tax bases, lower living standards, and disempower citizens--risks that mirror the 40-year stagnation of wages during the first in…
Query Circuits: Explaining How Language Models Answer User Prompts
Tung-Yu Wu, Fazl Barez
Explaining why a language model produces a particular output requires local, input-level explanations. Existing methods uncover global capability circuits (e.g., indirect object id…
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
Florent Draye, Abir Harrasse, Vedant Palit +8
Mechanistic interpretability seeks to understand how Large Language Models (LLMs) represent and process information. Recent approaches based on dictionary learning and transcoders…
U-shaped and Inverted-U Scaling behind Emergent Abilities of Large Language Models
Tung-Yu Wu, Pei-Yu Lo
Large language models (LLMs) have been shown to exhibit emergent abilities in some downstream tasks, where model performance stagnates at first and then improves sharply and unpred…