collaborators

10 papers

cs.CV2026

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models

Hongyu Zhang, Cheng Yan, Xiang Xia +1

Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full expert pool. Training-free expe…

cs.CV2026

DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models

Yongkang Zhou, Xiang Xia, Cheng Yan +2

Diffusion vision-language models (dVLMs) iteratively denoise masked responses while conditioning each denoising step on visual evidence, making visual conditioning a substantial re…

cs.AI2026

REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models

Xiang Xia, Cheng Yan, Yiming Zhang +3

Mixture-of-experts (MoE) models increase parameter capacity by activating only a small subset of experts for each token. This conditional-computation paradigm has enabled autoregre…

cs.AI2026

UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention

Cheng Yan, Guangyang Ye, Wuyang Zhang +5

While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can also exacerbate overthinking and…

cs.LG2026

DepCap: Adaptive Block-Wise Parallel Decoding for Efficient Diffusion LM Inference

Xiang Xia, Wuyang Zhang, Jiazheng Liu +2

Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive language generation due to their potential for parallel decoding and global refinement of…

cs.RO2026

FAVLA: A Force-Adaptive Fast-Slow VLA model for Contact-Rich Robotic Manipulation

Yao Li, Peiyuan Tang, Wuyang Zhang +7

Force/torque feedback can substantially improve Vision-Language-Action (VLA) models on contact-rich manipulation, but most existing approaches fuse all modalities at a single opera…