activity
20242026
collaborators

6 papers

cs.CL2026

A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results

Thai-Binh Nguyen, Katerina Zmolikova, Pingchuan Ma +3

We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a s…

cs.CL2026

VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation

Yancheng Wang, Osama Hanna, Ruiming Xie +11

Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features suc…

cs.SE2026

The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes

Redacted by arXiv

This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…

cs.CL2025

Overcoming Latency Bottlenecks in On-Device Speech Translation: A Cascaded Approach with Alignment-Based Streaming MT

Zeeshan Ahmed, Frank Seide, Niko Moritz +5

This paper tackles several challenges that arise when integrating Automatic Speech Recognition (ASR) and Machine Translation (MT) for real-time, on-device streaming speech translat…

cs.CL2025

Non-Monotonic Attention-based Read/Write Policy Learning for Simultaneous Translation

Zeeshan Ahmed, Frank Seide, Zhe Liu +6

Simultaneous or streaming machine translation generates translation while reading the input stream. These systems face a quality/latency trade-off, aiming to achieve high translati…

eess.AS2024

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition

Niko Moritz, Ruiming Xie, Yashesh Gaur +5

We propose the joint speech translation and recognition (JSTAR) model that leverages the fast-slow cascaded encoder architecture for simultaneous end-to-end automatic speech recogn…