3 papers
cs.CL2026
VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models
Luc Debaupte, Tyler Baumgartner, Brandon Tai +3
Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-only benchmark for evaluating…
cs.CL2026
Long-range Modeling and Processing of Multimodal Event Sequences
Jichu Li, Yilun Zhong, Zhiting Li +2
Temporal point processes (TPPs) have emerged as powerful tools for modeling asynchronous event sequences. While recent advances have extended TPPs to handle textual information, ex…
cs.CV2025
PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities
Jiajun Chen, Sai Cheng, Yutao Yuan +4
Multimodal models integrating natural language and visual information have substantially improved generalization of representation models. However, their effectiveness significantl…