activity
20232026
most citedThe Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts

3 citations · 5 across the 13 of their papers we have counts for

collaborators
Showing 2026Show all

6 papers · 1 filter

cs.SD2026

Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages

Haopeng Xie, Ismail Rasim Ulgen, Sofia Son +2

Prosody plays an important role in speech translation, conveying information such as emphasis, emotion, and intent beyond lexical content. However, despite recent progress in expre…

cs.CL2026

Multilingual Emotion Neurons in Large Audio-Language Models

Xiutian Zhao, Philipp Koehn, Björn Schuller +1

Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks,…

eess.AS2026

TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion

Zihan Zhang, Shreeram Suresh Chandra, Zongyang Du +5

Traditional emotional voice conversion (EVC) conditions generation on explicit target emotions like labels or references, defining the target affective state but omitting the direc…

eess.AS2026

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews +2

To preserve or not to preserve prosody is a central question in voice anonymization. Prosody conveys meaning and affect, yet is tightly coupled with speaker identity. Existing meth…

cs.CL2026

Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models

Xiutian Zhao, Ismail Rasim Ulgen, Philipp Koehn +2

Large audio-language models (LALMs) can produce expressive speech, yet reliable emotion control remains elusive: conversions often miss the target affect and may degrade linguistic…

cs.CV2026

FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation

Weiting Tan, Andy T. Liu, Ming Tu +3

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate…