4 papers
Perturbation-Based Uncertainty for Failure Detection in Vision-Language-Action Models
Yousung Lee, Dongsoo Har
Vision-Language-Action (VLA) models have shown strong performance in robotic manipulation, but reliable uncertainty quantification remains challenging, particularly under distribut…
Steering Sparse Autoencoder Latents to Control Dynamic Head Pruning in Vision Transformers (Student Abstract)
Yousung Lee, Dongsoo Har
Dynamic head pruning in Vision Transformers (ViTs) improves efficiency by removing redundant attention heads, but existing pruning policies are often difficult to interpret and con…
Circuit Fingerprints: How Answer Tokens Encode Their Geometrical Path
Andres Saurez, Neha Sengar, Dongsoo Har
Circuit discovery and activation steering in transformers have developed as separate research threads, yet both operate on the same representational space. Are they two views of th…
Why Linear Interpretability Works: Invariant Subspaces as a Result of Architectural Constraints
Andres Saurez, Yousung Lee, Dongsoo Har
Linear probes and sparse autoencoders consistently recover meaningful structure from transformer representations -- yet why should such simple methods succeed in deep, nonlinear sy…