collaborators

6 papers

cs.AI2026

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Haoyu Zhang, Zhipeng Li, Xiaoying Tang +2

Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce \textbf{Ex-Omni-2D}, an omn…

cs.AI2026

Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization

Haoyue Liu, Xiaoyu Ma, Ye Chen +2

Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding promising results. However,…

cs.LG2026

Diffusion-based learning framework for Constrained Nonconvex Optimization with Weighted Bootstrapped Refinement

Shutong Ding, Yimiao Zhou, Ke Hu +4

Recent advances in diffusion models show promising potential to accelerate nonconvex problem solving by leveraging their multimodality. However, most existing diffusion-based optim…

cs.CV2026

RAVE: Re-Allocating Visual Attention in Large Multimodal Models

Xi Leng, Xinhong Ma, Ziqiang Dong +4

Large multimodal models (LMMs) inherit the self-attention mechanism of pretrained language backbones, yet standard attention can exhibit suboptimal allocation, including cross-moda…

cs.CV2026

NLPrompt: Noise-Label Prompt Learning for Vision-Language Models

Bikang Pan, Qun Li, Xiaoying Tang +6

The emergence of vision-language foundation models, such as CLIP, has revolutionized image-text representation, enabling a broad range of applications via prompt learning. Despite…

cs.AI2025

FLEx: Personalized Federated Learning for Mixture-of-Experts LLMs via Expert Grafting

Fan Liu, Bikang Pan, Zhongyi Wang +4

Federated instruction tuning of large language models (LLMs) is challenged by significant data heterogeneity across clients, demanding robust personalization. The Mixture of Expert…