15 papers
MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models
Yuncheng Yang, Feiyang Ye, Shixian Luo +7
Vision-Language Models (VLMs) have achieved success using homogeneous Transformers to process multimedia data. Recent studies show that heterogeneous structures interleaving effici…
-Prediction Flow: Efficient Continuous Decoding for Masked Diffusion Language Models
Weitian Wang, Lianlei Shan, Shubham Rai +2
Masked diffusion language models (MDLMs) generate text by iteratively unmasking tokens, but their standard decoder reduces each step to a binary action: a position is either commit…
PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space
Bochen Yang, Lianlei Shan
Current Vision-Language-Action (VLA) models face a trade-off between efficient action generation and explicit deliberation. Directly decoding actions from vision-language backbone…
DLWM: Diverse Latent World Models for Efficient Multimodal Reasoning
David Huang, Lianlei Shan
Reasoning capabilities of multimodal large language models (MLLMs) have improved considerably in recent years. Existing approaches typically rely on explicit chain-of-thought or co…
Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action Models
Dianqiao Lei, Lianlei Shan
Existing Vision-Language-Action (VLA) models predominantly rely on explicit Chain-of-Thought (CoT) reasoning to bridge perception and action. While effective, this paradigm suffers…
MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action
Boyang Zhang, Lianlei Shan
Vision-Language-Action (VLA) policies remain brittle in long-horizon and high-uncertainty control, where one-pass action decoding provides limited inference-time deliberation. Expl…