From the 1 of 4 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Transcoders for Investigating Deception in Language Models
Darius Lim, Nathan Leow, Xin Wei Chia
The paper uses per‑layer transcoders to build attribution graphs that reveal internal features linked to deceptive outputs in a Qwen3‑4B language model, showing how deception can b…
cs.AI2026
Multi-Trait Subspace Steering to Reveal the Dark Side of Human-AI Interaction
Xin Wei Chia, Swee Liang Wong, Jonathan Pan
Recent incidents have highlighted alarming cases where human-AI interactions led to negative psychological outcomes, including mental health crises and even user harm. As LLMs serv…