5 citations · 5 across the 3 of their papers we have counts for
3 papers
CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model
Nhan Phan, Ilona Lähteenmäki, Anna von Zansen +4
Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies…
Advancing Audio Emotion and Intent Recognition with Large Pre-Trained Models and Bayesian Inference
Dejan Porjazovski, Yaroslav Getman, Tamás Grósz +1
Large pre-trained models are essential in paralinguistic systems, demonstrating effectiveness in tasks like emotion recognition and stuttering detection. In this paper, we employ l…
Topic Identification For Spontaneous Speech: Enriching Audio Features With Embedded Linguistic Information
Dejan Porjazovski, Tamás Grósz, Mikko Kurimo
Traditional topic identification solutions from audio rely on an automatic speech recognition system (ASR) to produce transcripts used as input to a text-based model. These approac…