4 papers · 1 filter
LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding
Guang Yang, Brian Siyuan Zheng, Victoria Ebert +1
We propose a novel pipeline, Legato 2, for extracting symbolic notation and semantic knowledge from images of sheet music. Legato 2 features the first large-scale neural model for…
Unleashing Video Language Models for Fine-grained HRCT Report Generation
Yingying Fang, Huichi Zhou, KinHei Lee +4
Generating precise diagnostic reports from High-Resolution Computed Tomography (HRCT) is critical for clinical workflow, yet it remains a formidable challenge due to the high patho…
Multimodal OCR: Parse Anything from Documents
Handong Zheng, Yumeng Li, Kaile Zhang +22
We present Multimodal OCR (MOCR), a document parsing paradigm that jointly parses text and graphics into unified textual representations. Unlike conventional OCR systems that focus…
LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR
Guang Yang, Victoria Ebert, Nazif Tamer +3
We propose Legato, a new end-to-end model for optical music recognition (OMR), a task of converting music score images to machine-readable documents. Legato is the first large-scal…