7 papers
Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders
Sergio Lanza, Jae Hee Lee, Stefan Wermter
Vision Language Models (VLMs) have demonstrated impressive performance in tasks requiring joint understanding of images and text, such as image captioning and Visual Question Answe…
Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR
Henri-Leon Kordt, Theresa Pekarek Rosin, Jae Hee Lee +1
Despite advances in large-scale Automatic Speech Recognition (ASR), disfluent speech remains challenging, as state-of-the-art systems are often optimized to omit disfluencies, lead…
Bias Leaves a Gradient Trail: Label-Free Bias Identification via Gradient Probes on Concept Decompositions
Thomas Vitry, Kieran Edgeworth, Stefan Wermter +1
Vision classifiers can exploit spurious correlations, achieving high in-distribution accuracy yet failing under distribution shift. Existing approaches to bias mitigation and analy…
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
Jeremy Herbst, Stefan Wermter, Jae Hee Lee
Mixture-of-Experts (MoE) architectures have become the dominant choice for scaling Large Language Models (LLMs), activating only a subset of parameters per token. While MoE archite…
Explaining, Verifying, and Aligning Semantic Hierarchies in Vision-Language Model Embeddings
Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein +9
Vision-language model (VLM) encoders such as CLIP enable strong retrieval and zero-shot classification in a shared image-text embedding space, yet the semantic organization of this…
Knowing the Facts but Choosing the Shortcut: Understanding How Large Language Models Compare Entities
Hans Hergen Lehmann, Jae Hee Lee, Steven Schockaert +1
Large Language Models (LLMs) are increasingly used for knowledge-based reasoning tasks, yet understanding when they rely on genuine knowledge versus superficial heuristics remains…