3 papers
cs.LG2026
Titans-as-a-Layer: Test-Time Memory for Conversational Speech Emotion Recognition
Daniel Chen, Qicong Hu, Yang Xiao +2
Speech emotion recognition (SER) is commonly formulated as utterance-level classification, although conversational emotion depends on a speaker's usual vocal range and the emotiona…
cs.CL2025
Reverb: Open-Source ASR and Diarization from Rev
Nishchal Bhandari, Danny Chen, Miguel Ãngel del RÃo Fernández +10
Today, we are open-sourcing our core speech recognition and diarization models for non-commercial use. We are releasing both a full production pipeline for developers as well as pa…
cs.CL2024
Style-agnostic evaluation of ASR using multiple reference transcripts
Quinten McNamara, Miguel Ãngel del RÃo Fernández, Nishchal Bhandari +4
Word error rate (WER) as a metric has a variety of limitations that have plagued the field of speech recognition. Evaluation datasets suffer from varying style, formality, and inhe…