2 papers
cs.CL2026
The Neutral Mask: How RLHF Provides Shallow Alignment while Leaving Partisan Structure Intact in a Large Language Model
Wendy K. Tam
The ambition behind alignment training is to make large language models safe and useful. The primary mechanism, reinforcement learning from human feedback (RLHF), shapes the behavi…
cs.CL2026
The Amplifying Mirror: Locating and Steering the Partisan Direction inside a Large Language Model
Wendy K. Tam
Large language models are rapicly replacing search engines as the primary interface between people and information. Unlike search engines, which retrieve existing content, LLMs gen…