4 papers
ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
Juze Zhang, Changan Chen, Xin Chen +5
Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human behavior as a translation task c…
MAEB: Massive Audio Embedding Benchmark
Adnan El Assadi, Isaac Chung, Chenghao Xiao +15
We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasonin…
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
Potsawee Manakul, Woody Haosheng Gan, Michael J. Ryan +5
Current speech evaluation suffers from two critical limitations: the need and difficulty of designing specialized systems targeting individual audio characteristics, and poor corre…
Multi-Stage Speaker Diarization for Noisy Classrooms
Ali Sartaz Khan, Tolulope Ogunremi, Ahmed Adel Attia +1
Speaker diarization, the process of identifying "who spoke when" in audio recordings, is essential for understanding classroom dynamics. However, classroom settings present distinc…