14 papers
Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions
Quyen Tran, Hai Nguyen, Quan Dao +4
Analytic Continual Learning (ACL) offers a computationally efficient alternative to gradient-based approaches. Recent ACL methods are based on Recursive Least Squares (RLS) and hav…
Few-Step Diffusion Language Models via Trajectory Self-Distillation
Tunyu Zhang, Xinxi Zhang, Ligong Han +9
Diffusion large language models (DLLMs) have emerged as powerful generative models with the promise of fast text generation through parallel decoding. However, realizing this poten…
Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers
Anh Nguyen, Ngan Nguyen, Duc Vu +11
Modern one-step diffusion models achieve impressive quality through distribution-based timestep distillation. Yet, they rely on a critical assumption: Teacher and Student must inha…
Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
Gerasimos Chatzoudis, Zhuowei Li, Gemma E. Moran +2
Sparse Autoencoders (SAEs) are increasingly used to interpret foundation models, but their role as an actionable intervention space remains less understood, especially in vision. W…
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
Gerasimos Chatzoudis, Konstantinos D. Polyzos, Zhuowei Li +4
Understanding the internal activations of Vision Transformers (ViTs) is critical for building interpretable and trustworthy models. While Sparse Autoencoders (SAEs) have been used…
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
Tunyu Zhang, Haizhou Shi, Yibin Wang +9
While Large Language Models (LLMs) have demonstrated impressive capabilities, their output quality remains inconsistent across various application scenarios, making it difficult to…