5 papers
Assessing Factual Music Comprehension in Large Audio Language Models
Daniel Chenyu Lin, Michael Freeman, John Thickstun
Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio. In this paper, we (1) provide empiri…
Music Transcription with (Almost) No Supervision
Saebyeol Shin, Chao Wan, Zhenzhen Liu +4
Competitive music transcription models require large amounts of paired audio-score data, which is scarce due to collection costs, alignment difficulty, and copyright restrictions.…
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
Haowei Fu, Bo Ni, Han Xu +3
Retrieval-Augmented Generation (RAG) and Supervised Finetuning (SFT) have become the predominant paradigms for equipping Large Language Models (LLMs) with external knowledge for di…
A Generative Foundation Model for Chest Radiography
Yuanfeng Ji, Dan Lin, Xiyue Wang +12
The scarcity of well-annotated diverse medical images is a major hurdle for developing reliable AI models in healthcare. Substantial technical advances have been made in generative…
olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
Jake Poznanski, Aman Rangapur, Jon Borchardt +6
PDF documents have the potential to provide trillions of novel, high-quality tokens for training language models. However, these documents come in a diversity of types with differi…