activity
20242026
collaborators

6 papers

cs.AI2026

Voxtral TTS

Mistral-AI, :, Alexander H. Liu +186

We introduce Voxtral TTS, an expressive multilingual text-to-speech model that generates natural speech from as little as 3 seconds of reference audio. Voxtral TTS adopts a hybrid…

cs.AI2026

Voxtral Realtime

Mistral-AI, :, Alexander H. Liu +166

We introduce Voxtral Realtime, a natively streaming automatic speech recognition model that matches offline transcription quality at sub-second latency. Unlike approaches that adap…

cs.LG2026

On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks

Yannic Neuhaus, Nicolas Flammarion, Matthias Hein +1

Integrating reasoning in large language models and large vision-language models has recently led to significant improvement of their capabilities. However, the generalization of re…

cs.CV2025

RePOPE: Impact of Annotation Errors on the POPE Benchmark

Yannic Neuhaus, Matthias Hein

Since data annotation is costly, benchmark datasets often incorporate labels from established image datasets. In this work, we assess the impact of label errors in MSCOCO on the fr…

cs.CV2025

DASH: Detection and Assessment of Systematic Hallucinations of VLMs

Maximilian Augustin, Yannic Neuhaus, Matthias Hein

Vision-language models (VLMs) are prone to object hallucinations, where they erroneously indicate the presenceof certain objects in an image. Existing benchmarks quantify hallucina…

cs.CV2024

DiG-IN: Diffusion Guidance for Investigating Networks -- Uncovering Classifier Differences Neuron Visualisations and Visual Counterfactual Explanations

Maximilian Augustin, Yannic Neuhaus, Matthias Hein

While deep learning has led to huge progress in complex image classification tasks like ImageNet, unexpected failure modes, e.g. via spurious features, call into question how relia…