collaborators

8 papers

cs.CL2026

SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment

Tianyu Dong, Yangyang Liu, Jiang Zhou +9

Sparse Mixture-of-Experts (MoE) architectures have emerged as an increasingly influential paradigm as they offer a strategic balance between parameter scalability and computational…

cs.LG2026

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference

Bo Li, Chuan Wu, Shaolin Zhu

Mixture-of-Experts Multimodal Large Language Models (MoE MLLMs) suffer from a significant efficiency bottleneck during Expert Parallelism (EP) inference due to the straggler effect…

cs.LG2026

TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts

Jiangyang He, Shaolin Zhu, Deyi Xiong

Mixture-of-Experts large language models (LLMs) scale efficiently through sparse activation, yet their deployment is fundamentally constrained by the large static parameter footpri…

cs.CL2026

Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs

Bo Li, Tianyu Dong, Shaolin Zhu +1

Large Language Models (LLMs) have shown great promise in multilingual machine translation (MT), even with limited bilingual supervision. However, fine-tuning LLMs with parallel cor…

cs.CV2026

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation

Bo Li, Ronghao Chen, Ningyuan Deng +3

Translating text embedded in Web images is crucial for improving content accessibility and cross-lingual information retrieval, particularly within social media and e-commerce doma…

cs.CL2026

MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation

Bo Li, Ningyuan Deng, Tianyu Dong +3

Multimodal large language models (MLLMs) have shown impressive capabilities, yet they often struggle to effectively capture the fine-grained textual information within images cruci…