12 papers
From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons
Xiangyu Ma, Teng Xiao, Zuchao Li +1
Diffusion models promise efficient parallel text generation but rely on bidirectional attention, creating a structural mismatch with pre-trained Autoregressive (AR) models. This in…
SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away
Jiajia Li, Jiliang Hu, Ziyi Pan +4
Recently, there have been significant advancements in music generation. However, existing models primarily focus on creating modern pop songs, making it challenging to produce anci…
OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
Teng Xiao, Zuchao Li, Lefei Zhang
Recent advances in multimodal large language models (LLMs) have led to significant progress in understanding, generation, and retrieval tasks. However, current solutions often trea…
Model Hemorrhage and the Robustness Limits of Large Language Models
Ziyang Ma, Zuchao Li, Lefei Zhang +4
Large language models (LLMs) demonstrate strong performance across natural language processing tasks, yet undergo significant performance degradation when modified for deployment t…
Label Drop for Multi-Aspect Relation Modeling in Universal Information Extraction
Lu Yang, Jiajia Li, En Ci +3
Universal Information Extraction (UIE) has garnered significant attention due to its ability to address model explosion problems effectively. Extractive UIE can achieve strong perf…
NOTA: Multimodal Music Notation Understanding for Visual Large Language Model
Mingni Tang, Jiajia Li, Lu Yang +5
Symbolic music is represented in two distinct forms: two-dimensional, visually intuitive score images, and one-dimensional, standardized text annotation sequences. While large lang…