2 papers
cs.RO2026
-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs
Siting Wang, Xiaofeng Wang, Zheng Zhu +7
Flow-based vision-language-action (VLA) models excel in embodied control but suffer from intractable likelihoods during multi-step sampling, hindering online reinforcement learning…
cs.CV2026
CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning
Chunlei Meng, Guanhong Huang, Rong Fu +3
Multimodal learning aims to capture both shared and private information from multiple modalities. However, existing methods that project all modalities into a single latent space f…