4 papers
DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Haoxiang Sun, Lizhen Xu, Bing Zhao +5
Reinforcement Learning with Verifiable Rewards (RLVR) has been shown effective in enhancing the visual reflection and reasoning capabilities of Large Multimodal Models (LMMs). Howe…
SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning
Xu Pan, Zhenglin Wan, Xingrui Yu +6
Vision-Language-Action (VLA) models exhibit strong generalization in robotic manipulation, yet reinforcement learning (RL) fine-tuning often degrades robustness under spatial distr…
Beyond flattening: a geometrically principled positional encoding for vision transformers with Weierstrass elliptic functions
Zhihang Xin, Xitong Hu, Rui Wang
Vision Transformers have demonstrated remarkable success in computer vision tasks, yet their reliance on learnable one-dimensional positional embeddings fundamentally disrupts the…
Scaling Laws of Motion Forecasting and Planning -- Technical Report
Mustafa Baniodeh, Kratarth Goel, Scott Ettinger +14
We study the empirical scaling laws of a family of encoder-decoder autoregressive transformer models on the task of joint motion forecasting and planning in the autonomous driving…