6 papers
Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models
Hongyu Zhang, Cheng Yan, Xiang Xia +1
Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full expert pool. Training-free expe…
DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models
Yongkang Zhou, Xiang Xia, Cheng Yan +2
Diffusion vision-language models (dVLMs) iteratively denoise masked responses while conditioning each denoising step on visual evidence, making visual conditioning a substantial re…
UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention
Cheng Yan, Guangyang Ye, Wuyang Zhang +5
While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can also exacerbate overthinking and…
Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer
Yuze Li, Dong Gong, Xiao Cao +6
Motion transfer has emerged as a promising direction for controllable video generation, yet existing methods largely focus on single-object scenarios and struggle when multiple obj…
Learning Visual Proxy for Compositional Zero-Shot Learning
Shiyu Zhang, Cheng Yan, Yang Liu +3
Compositional Zero-Shot Learning (CZSL) aims to recognize novel attribute-object compositions by leveraging knowledge from seen compositions. Current methods align textual prototyp…
ROS-SAM: High-Quality Interactive Segmentation for Remote Sensing Moving Object
Zhe Shan, Yang Liu, Lei Zhou +3
The availability of large-scale remote sensing video data underscores the importance of high-quality interactive segmentation. However, challenges such as small object sizes, ambig…