activity
20162026
most citedMERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

28 citations · 169 across the 76 of their papers we have counts for

collaborators
Showing eess.ASShow all

32 papers · 1 filter

eess.AS2025

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

Yinghao Ma, Siyou Li, Juntao Yu +2

Recent advances in audio-text large language models (LLMs) have opened new possibilities for music understanding and generation. However, existing benchmarks are limited in scope,…

eess.AS2025

Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss

Jiawen Huang, Felipe Sousa, Emir Demirel +2

Automatic Lyrics Transcription (ALT) aims to recognize lyrics from singing voices, similar to Automatic Speech Recognition (ASR) for spoken language, but faces added complexity due…

eess.AS2025

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

Huan Zhang, Jinhua Liang, Huy Phan +2

Evaluating generative models remains a fundamental challenge, particularly when the goal is to reflect human preferences. In this paper, we use music generation as a case study to…

eess.AS2025★ 1 cited

YuE: Scaling Open Foundation Models for Long-Form Music Generation

Ruibin Yuan, Hanfeng Lin, Shuyue Guo +55

We tackle the task of long-form music generation--particularly the challenging \textbf{lyrics-to-song} problem--by introducing YuE, a family of open foundation models based on the…

eess.AS2024

YourMT3+: Multi-instrument Music Transcription with Enhanced Transformer Architectures and Cross-dataset Stem Augmentation

Sungkyun Chang, Emmanouil Benetos, Holger Kirchhoff +1

Multi-instrument music transcription aims to convert polyphonic music recordings into musical scores assigned to each instrument. This task is challenging for modeling as it requir…

eess.AS2024

Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model

Jiawen Huang, Emmanouil Benetos

Multilingual automatic lyrics transcription (ALT) is a challenging task due to the limited availability of labelled data and the challenges introduced by singing, compared to multi…