Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
Eric Onyame, Runtao Zhou, Kowshik Thopalli +2
Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. However, its reliability remains lar…
cs.CL2025
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
Xinyuan Yan, Shusen Liu, Kowshik Thopalli +1
Sparse autoencoders (SAEs) have emerged as a powerful tool for uncovering interpretable features in large language models (LLMs) through the sparse directions they learn. However,…