activity
20242026
collaborators

11 papers

cs.CV2026

Identifying the Unknown: Prompt-Free Open Vocabulary Anomaly Recognition for Robot-Object Interaction

Philipp Allgeuer, Jan-Gerrit Habekost, Stefan Wermter

Robots operating in real-world environments must in general be able to recognize previously unseen objects. As robotic systems move toward open-world autonomy, there is a growing,…

cs.CV2026

Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders

Sergio Lanza, Jae Hee Lee, Stefan Wermter

Vision Language Models (VLMs) have demonstrated impressive performance in tasks requiring joint understanding of images and text, such as image captioning and Visual Question Answe…

cs.SD2026

A Generalized Formalism of Auto-Regressive Decoding for Speech Processing

Julia Gachot, Philipp Allgeuer, Marie S. Bauer +1

In speech processing, most state-of-the-art sequence prediction models rely on auto-regressive (AR) strategies to generate output sequences based on the raw predictions of the mode…

cs.CL2026

MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition

Theresa Pekarek Rosin, Matthias Kerzel, Stefan Wermter

Modern Automatic Speech Recognition (ASR) systems have made remarkable progress on standard benchmarks, yet performance gaps have emerged under real-world distribution shifts, caus…

cs.CL2026

Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR

Henri-Leon Kordt, Theresa Pekarek Rosin, Jae Hee Lee +1

Despite advances in large-scale Automatic Speech Recognition (ASR), disfluent speech remains challenging, as state-of-the-art systems are often optimized to omit disfluencies, lead…

cs.CL2026

The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level

Jeremy Herbst, Stefan Wermter, Jae Hee Lee

Mixture-of-Experts (MoE) architectures have become the dominant choice for scaling Large Language Models (LLMs), activating only a subset of parameters per token. While MoE archite…