3 papers
cs.SD2026
Gen-SER: When the generative model meets speech emotion recognition
Taihui Wang, Jinzheng Zhao, Rilin Chen +3
Speech emotion recognition (SER) is crucial in speech understanding and generation. Most approaches are based on either classification models or large language models. Different fr…
cs.SD2025
Vox-Evaluator: Enhancing Stability and Fidelity for Zero-shot TTS with A Multi-Level Evaluator
Hualei Wang, Na Li, Chuke Wang +3
Recent advances in zero-shot text-to-speech (TTS), driven by language models, diffusion models and masked generation, have achieved impressive naturalness in speech synthesis. Neve…
cs.AI2025
StarBench: A Turn-Based RPG Benchmark for Agentic Multimodal Decision-Making and Information Seeking
Haoran Zhang, Chenhao Zhu, Sicong Guo +3
Human players do more than press buttons: they ground what they see on screen into precise keyboard-mouse actions and, when stuck, they seek information before trying again. We ask…