From the 1 of 31 linked papers with an AI index.
31 papers
The Role of Disfluencies in Speech Translation
Maike Züfle, Maria Teleki, Fabian Retkowski +5
Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts rather than translate them.…
When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
Wendi Yu, Lianhao Zhou, Xiangjue Dong +6
LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with perform…
Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
Tian Liu, Anwesha Basu, James Caverlee +1
The paper introduces a training-free post-hoc correction framework that uses large multimodal models to improve few-shot expert models for visual species recognition, boosting accu…
Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective
Tian Liu, Anwesha Basu, James Caverlee +1
Semi-supervised few-shot learning (SSFSL) resembles real-world applications such as auto-annotation, as it aims to learn a model from a few labeled and abundant unlabeled task-spec…
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
Shree Harsha Bokkahalli Satish, Christoph Minixhofer, Maria Teleki +5
Speech Large Language Models (SpeechLLMs) process spoken input directly, retaining cues such as accent and perceived gender that were previously removed in cascaded pipelines. This…
Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation
Xingyu Su, Jacob Helwig, Shubham Parashar +6
We study the transformation of autoregressive models (ARLMs) into diffusion language models (DLMs). Rather than pretraining from scratch, prior work replaces the causal attention i…