7 papers
UAF: A Unified Audio Front-end LLM for Full-Duplex Speech Interaction
Yadong Li, Guoxin Wu, Haiping Hou +1
Full-duplex speech interaction, as the most natural and intuitive mode of human communication, is driving artificial intelligence toward more human-like conversational systems. Tra…
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
Wei Song, Yuran Wang, Zijia Song +6
The differing representation spaces required for visual understanding and generation pose a challenge in unifying them within the autoregressive paradigm of large language models.…
CrossHGL: A Text-Free Foundation Model for Cross-Domain Heterogeneous Graph Learning
Xuanze Chen, Jiajun Zhou, Yadong Li +2
Heterogeneous graph representation learning (HGRL) is essential for modeling complex systems with diverse node and edge types. However, most existing methods are limited to closed-…
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
Song Chen, Xinyu Guo, Yadong Li +10
Multimodal large language models (MLLMs) have shown impressive capabilities across various domains, excelling in processing and understanding information from multiple modalities.…
Baichuan-Omni-1.5 Technical Report
Yadong Li, Jun Liu, Tao Zhang +89
We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve f…
Baichuan-Omni Technical Report
Yadong Li, Haoze Sun, Mingan Lin +23
The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpa…