Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Reducing Political Manipulation with Consistency Training
Long Phan, Devin Kim, Alexander Pan +3
Large language models (LLMs) exhibit systematic political bias across a variety of sensitive contexts. We find that LLMs handle counterpart topics from opposing political sides asy…
cs.CL2026
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
Alexander Pan, Lijie Chen, Jacob Steinhardt
Top-down transparency typically analyzes language model activations using probes with scalar or single-token outputs, limiting the range of behaviors that can be captured. To allev…
cs.CL2025
Context Is Not Comprehension
Alex Pan, Mary-Anne Williams
The dominant way of judging Large Language Models (LLMs) has been to ask how well they can recall explicit facts from very long inputs. While today's best models achieve near perfe…