2 papers
cs.SD2026
TargetSEC: Plug-and-Play In-the-Wild Speech Emotion Conversion via Arousal-Conditioned Latent Style Diffusion
Constantin Alexander Auga
Speech Emotion Conversion (SEC) aims to transform the emotion of a source utterance into a target emotion while preserving content and speaker identity. SEC on in-the-wild data is…
cs.CV2026
Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models
Weixing Wang, Liudvikas Zekas, Anton Hackl +5
Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evaluation protocols assess these…