4 papers
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis
Run Luo, Ting-En Lin, Haonan Zhang +10
Recent advancements in omnimodal learning have significantly improved understanding and generation across images, text, and speech, yet these developments remain predominantly conf…
IP-MOT: Instance Prompt Learning for Cross-Domain Multi-Object Tracking
Run Luo, Zikai Song, Longze Chen +3
Multi-Object Tracking (MOT) aims to associate multiple objects across video frames and is a challenging vision task due to inherent complexities in the tracking environment. Most e…
Ruler: A Model-Agnostic Method to Control Generated Length for Large Language Models
Jiaming Li, Lei Zhang, Yunshui Li +5
The instruction-following ability of large language models enables humans to interact with AI agents in a natural way. However, when required to generate responses of a specific le…
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
Run Luo, Haonan Zhang, Longze Chen +13
The development of Multimodal Large Language Models (MLLMs) has seen significant advancements with increasing demands in various fields (e.g., multimodal agents, embodied intellige…