5 papers
Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
Zihao Wang, Ruibin Yuan, Ziqi Geng +7
Automated singing assessment is crucial for education and entertainment. However, existing systems face two fundamental limitations: reliance on reference tracks, which stifles cre…
SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation
Chenyang Le, Bing Han, Jinshun Li +2
Simultaneous Speech Translation (SimulST) enables real-time cross-lingual communication by jointly optimizing speech recognition and machine translation under strict latency constr…
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
Song Chen, Xinyu Guo, Yadong Li +10
Multimodal large language models (MLLMs) have shown impressive capabilities across various domains, excelling in processing and understanding information from multiple modalities.…
Baichuan-Omni-1.5 Technical Report
Yadong Li, Jun Liu, Tao Zhang +89
We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve f…
Baichuan-Omni Technical Report
Yadong Li, Haoze Sun, Mingan Lin +23
The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpa…