5 papers
NewtonGS: Physics-Structured Object-Level Neural Newtonian Dynamics for Gaussian Scene Animation
Lianlei Shan, Feiyang Ye, Yan Chen +1
Animating objects in a static 3D Gaussian scene requires an explicit object-level dynamic state and a controllable model of object motion. Existing dynamic Gaussian methods primari…
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
Lei Xiao, Jifeng Li, Juntao Gao +6
Vision-Language-Action (VLA) models have shown remarkable progress in embodied tasks recently, but most methods process visual observations independently at each timestep. This his…
HBVLA: Pushing 1-Bit Post-Training Quantization for Vision-Language-Action Models
Xin Yan, Zhenglin Wan, Feiyang Ye +4
Vision-Language-Action (VLA) models enable instruction-following embodied control, but their large compute and memory footprints hinder deployment on resource-constrained robots an…
Faster and Better Alignment for Flow Matching Models via Step-aware Advantages
Zhixiong Yue, Zixuan Ni, Feiyang Ye +4
Recent advances in flow matching models, particularly with reinforcement learning (RL), have significantly enhanced human preference alignment in few-step text-to-image generators.…
Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
Juntao Gao, Feiyang Ye, Jing Zhang +1
Vision-Language-Action (VLA) models have emerged as a powerful paradigm in Embodied AI. However, the significant computational overhead of processing redundant visual tokens remain…