9 papers
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference
Bo Li, Chuan Wu, Shaolin Zhu
Mixture-of-Experts Multimodal Large Language Models (MoE MLLMs) suffer from a significant efficiency bottleneck during Expert Parallelism (EP) inference due to the straggler effect…
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Xiang An, Yin Xie, Feilong Tang +27
We introduce LLaVA-OneVision-2 (LLaVA-OV-2), the most capable vision-language model in the LLaVA-OneVision series to date, achieving superior performance across a broad range of mu…
Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
Bo Li, Tianyu Dong, Shaolin Zhu +1
Large Language Models (LLMs) have shown great promise in multilingual machine translation (MT), even with limited bilingual supervision. However, fine-tuning LLMs with parallel cor…
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
Bo Li, Ronghao Chen, Ningyuan Deng +3
Translating text embedded in Web images is crucial for improving content accessibility and cross-lingual information retrieval, particularly within social media and e-commerce doma…
DisastRAG: A Multi-Source Disaster Information Integration and Access System Based on Retrieval-Augmented Large Language Models
Bo Li, Zhitong Chen, Kai Yin +3
Effective disaster management requires rapid access to information distributed across structured operational records, unstructured institutional documents, and dynamic external sou…
Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection
Sa Zhu, Wanqian Zhang, Lin Wang +3
Open-Vocabulary Temporal Action Detection (OV-TAD) aims to localize and classify action segments of unseen categories in untrimmed videos, where effective alignment between action…