Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Certified Circuits: Stability Guarantees for Mechanistic Circuits
Alaa Anani, Tobias Lorenz, Bernt Schiele +2
Understanding how neural networks arrive at their predictions is essential for debugging, auditing, and deployment. Mechanistic interpretability pursues this goal by identifying ci…
cs.AI2026
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
Nina Żukowska, Wolfgang Stammer, Bernt Schiele +1
Transparency of neural networks' internal reasoning is at the heart of interpretability research, adding to trust, safety, and understanding of these models. The field of mechanist…
cs.AI2025
LayerCake: Token-Aware Contrastive Decoding within Large Language Model Layers
Jingze Zhu, Yongliang Wu, Wenbo Zhu +7
Large language models (LLMs) excel at natural language understanding and generation but remain vulnerable to factual errors, limiting their reliability in knowledge-intensive tasks…