2 papers
cs.CV2024
HRVDA: High-Resolution Visual Document Assistant
Chaohu Liu, Kun Yin, Haoyu Cao +6
Leveraging vast training data, multimodal large language models (MLLMs) have demonstrated formidable general visual comprehension capabilities and achieved remarkable performance a…
cs.SD2023
AudioFormer: Audio Transformer learns audio feature representations from discrete acoustic codes
Zhaohui Li, Haitao Wang, Xinghua Jiang
We propose a method named AudioFormer,which learns audio feature representations through the acquisition of discrete acoustic codes and subsequently fine-tunes them for audio class…