2 papers
cs.SD2026
Towards Realistic Synthetic Data for Automatic Drum Transcription
Pierfrancesco Melucci, Paolo Merialdo, Taketo Akama
Deep learning models define the state-of-the-art in Automatic Drum Transcription (ADT), yet their performance is contingent upon large-scale, paired audio-MIDI datasets, which are…
cs.CL2025
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
Francesco Verdini, Pierfrancesco Melucci, Stefano Perna +9
The remarkable performance achieved by Large Language Models (LLM) has driven research efforts to leverage them for a wide range of tasks and input modalities. In speech-to-text (S…