2 papers
cs.SD2026
Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
Congren Dai, Yue Yang, Krinos Li +12
Understanding complete musical scores entails integrated reasoning over pitch, rhythm, harmony, and large-scale structure, yet the ability of Large Language Models and Vision--Lang…
cs.CL2025
How Much of Your Data Can Suck? Thresholds for Domain Performance and Emergent Misalignment in LLMs
Jian Ouyang, Arman T, Ge Jin
This paper investigates the impact of incorrect data on the performance and safety of large language models (LLMs), specifically gpt-4o, during supervised fine-tuning (SFT). Althou…