4 papers
S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling
Dingkun Zhou, Shuchang Pan, Jiachen Lian +11
Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing this perceptual pathway is critical for b…
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
Dingkun Zhou, Krish Patel, Ajay Kankipati +14
Emotions conveyed through voice and face shape engagement and context in human AI interaction. Despite rapid progress in omni modal large language models, the holistic evaluation o…
EDMFormer: Genre-Specific Self-Supervised Learning for Music Structure Segmentation
Sahal Sajeer, Krish Patel, Oscar Chung +1
Music structure segmentation is a key task in audio analysis, but existing models perform poorly on Electronic Dance Music (EDM). This problem exists because most approaches rely o…
LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
Zongli Ye, Jiachen Lian, Akshaj Gupta +18
Phonetic speech transcription is crucial for fine-grained linguistic analysis and downstream speech applications. While Connectionist Temporal Classification (CTC) is a widely used…