13 papers
MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models
Hao Shi, Weiye Li, Bin Xie +6
Temporal modeling is essential for robotic manipulation, as effective control requires both memory of past interactions and imagination of future states. However, most VLA models r…
Towards World Models in Biomedical Research
Guangyu Wang, Jingkun Yue, Siqi Zhang +19
A central goal of biomedicine is to understand, predict and ultimately control the dynamic mechanisms by which biological systems respond to perturbations, disease progression and…
Linearizing Vision Transformer with Test-Time Training
Yining Li, Dongchen Han, Zeyu Liu +3
While linear-complexity attention mechanisms offer a promising alternative to Softmax attention for overcoming the quadratic bottleneck, training such models from scratch remains p…
Bridging the RGB-IR Gap: Consensus and Discrepancy Modeling for Text-Guided Multispectral Detection
Jiaqi Wu, Zhen Wang, Enhao Huang +5
Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust perception. However, notable limit…
AdaGen: Learning Adaptive Policy for Image Synthesis
Zanlin Ni, Yulin Wang, Yeguo Hua +5
Recent advances in image synthesis have been propelled by powerful generative models, such as Masked Generative Transformers (MaskGIT), autoregressive models, diffusion models, and…
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
Renping Zhou, Zanlin Ni, Tianyi Chen +6
Recently, Masked Diffusion Models (MDMs) have shown promising potential across vision, language, and cross-modal generation. However, a notable discrepancy exists between their tra…