8 citations · 13 across the 6 of their papers we have counts for
6 papers
Multimodal Emotion Regression with Multi-Objective Optimization and VAD-Aware Audio Modeling for the 10th ABAW EMI Track
Jiawen Huang, Chenxi Huang, Zhuofan Wen +7
We participated in the 10th ABAW Challenge, focusing on the Emotional Mimicry Intensity (EMI) Estimation track on the Hume-Vidmimic2 dataset. This task aims to predict six continuo…
Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss
Jiawen Huang, Felipe Sousa, Emir Demirel +2
Automatic Lyrics Transcription (ALT) aims to recognize lyrics from singing voices, similar to Automatic Speech Recognition (ASR) for spoken language, but faces added complexity due…
Foundation Models for Music: A Survey
Yinghao Ma, Anders Øland, Anton Ragni +39
In recent years, foundation models (FMs) such as large language models (LLMs) and latent diffusion models (LDMs) have profoundly impacted diverse sectors, including music. This com…
Improving Lyrics Alignment through Joint Pitch Detection
Jiawen Huang, Emmanouil Benetos, Sebastian Ewert
In recent years, the accuracy of automatic lyrics alignment methods has increased considerably. Yet, many current approaches employ frameworks designed for automatic speech recogni…
Modeling the Compatibility of Stem Tracks to Generate Music Mashups
Jiawen Huang, Ju-Chiang Wang, Jordan B. L. Smith +2
A music mashup combines audio elements from two or more songs to create a new work. To reduce the time and effort required to make them, researchers have developed algorithms that…
Score-informed Networks for Music Performance Assessment
Jiawen Huang, Yun-Ning Hung, Ashis Pati +2
The assessment of music performances in most cases takes into account the underlying musical score being performed. While there have been several automatic approaches for objective…