3 papers
cs.LG2026
Circuit Fingerprints: How Answer Tokens Encode Their Geometrical Path
Andres Saurez, Neha Sengar, Dongsoo Har
Circuit discovery and activation steering in transformers have developed as separate research threads, yet both operate on the same representational space. Are they two views of th…
cs.LG2026
Why Linear Interpretability Works: Invariant Subspaces as a Result of Architectural Constraints
Andres Saurez, Yousung Lee, Dongsoo Har
Linear probes and sparse autoencoders consistently recover meaningful structure from transformer representations -- yet why should such simple methods succeed in deep, nonlinear sy…
cs.CL2025
Continuous Adversarial Text Representation Learning for Affective Recognition
Seungah Son, Andrez Saurez, Dongsoo Har
While pre-trained language models excel at semantic understanding, they often struggle to capture nuanced affective information critical for affective recognition tasks. To address…