1 citations · 1 across the 3 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
NOTA: Multimodal Music Notation Understanding for Visual Large Language Model
Mingni Tang, Jiajia Li, Lu Yang +5
Symbolic music is represented in two distinct forms: two-dimensional, visually intuitive score images, and one-dimensional, standardized text annotation sequences. While large lang…
cs.CV2024
Multi-modal Auto-regressive Modeling via Visual Words
Tianshuo Peng, Zuchao Li, Lefei Zhang +3
Large Language Models (LLMs), benefiting from the auto-regressive modelling approach performed on massive unannotated texts corpora, demonstrates powerful perceptual and reasoning…