2 papers
cs.CL2025
CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition
Sina J. Semnani, Han Zhang, Xinyan He +2
Accurate text recognition for historical documents can greatly advance the study and preservation of cultural heritage. Existing vision-language models (VLMs), however, are designe…
cs.SD2023
InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models
Bing Han, Junyu Dai, Weituo Hao +6
Music editing primarily entails the modification of instrument tracks or remixing in the whole, which offers a novel reinterpretation of the original piece through a series of oper…