6 papers
Action with Visual Primitives
Weilong Guo, Yuchen Wang, Renping Zhou +5
Vision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robotic manipulation. A common design in current architectures maps language instructions an…
MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models
Hao Shi, Weiye Li, Bin Xie +6
Temporal modeling is essential for robotic manipulation, as effective control requires both memory of past interactions and imagination of future states. However, most VLA models r…
AdaGen: Learning Adaptive Policy for Image Synthesis
Zanlin Ni, Yulin Wang, Yeguo Hua +5
Recent advances in image synthesis have been propelled by powerful generative models, such as Masked Generative Transformers (MaskGIT), autoregressive models, diffusion models, and…
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
Renping Zhou, Zanlin Ni, Tianyi Chen +6
Recently, Masked Diffusion Models (MDMs) have shown promising potential across vision, language, and cross-modal generation. However, a notable discrepancy exists between their tra…
4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models
Wanhua Li, Renping Zhou, Jiawei Zhou +5
Learning 4D language fields to enable time-sensitive, open-ended language queries in dynamic scenes is essential for many real-world applications. While LangSplat successfully grou…
ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis
Zanlin Ni, Yulin Wang, Renping Zhou +5
Recently, token-based generation have demonstrated their effectiveness in image synthesis. As a representative example, non-autoregressive Transformers (NATs) can generate decent-q…