1 paper
Morteza Rohanian, Julian Hough, Matthew Purver
We present two multimodal fusion-based deep learning models that consume ASR transcribed speech and acoustic data simultaneously to classify whether a speaker in a structured diagn…