6 papers · 1 filter
Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models
Yixiang Liu, Zhongxing Xu, Zhonghua Wang +1
Diffusion vision-language models generate answers through iterative refinement, exposing intermediate answer trajectories that can be inspected and controlled at inference time. Ho…
Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection
Haoyue Liu, Xiaoyu Ma, Ye Chen +2
Reinforcement learning over a frozen reasoner has become a common recipe for teaching a policy which external tools to invoke. We show that this recipe becomes structurally mismatc…
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Haoyu Zhang, Zhipeng Li, Xiaoying Tang +2
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, but a spoken answer still leaves the agent visually absent. We introduce \textbf{Ex-Omni-…
Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval
Haoyue Liu, Ye Chen, Zhichao Wang +1
Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules into contrastive fine-tuning. However, the…
Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization
Haoyue Liu, Xiaoyu Ma, Ye Chen +2
Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding promising results. However,…
FLEx: Personalized Federated Learning for Mixture-of-Experts LLMs via Expert Grafting
Fan Liu, Bikang Pan, Zhongyi Wang +4
Federated instruction tuning of large language models (LLMs) is challenged by significant data heterogeneity across clients, demanding robust personalization. The Mixture of Expert…