9 papers
Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
Zijian Song, Xiaoxin Lin, Tao Pu +3
Recent progress in robotics and embodied AI is largely driven by Large Multimodal Models (LMMs). However, a key challenge remains underexplored: how can we advance LMMs to discover…
Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
Dapeng Zhang, Zhenlong Yuan, Zhangquan Chen +5
Vision-Language-Action (VLA) models have recently shown strong decision-making capabilities in autonomous driving. However, existing VLAs often struggle with achieving efficient in…
Pure Vision Language Action (VLA) Models: A Comprehensive Survey
Dapeng Zhang, Jing Sun, Chenghui Hu +5
The emergence of Vision Language Action (VLA) models marks a paradigm shift from traditional policy-based control to generalized robotics, reframing Vision Language Models (VLMs) f…
EMPOWER: Evolutionary Medical Prompt Optimization With Reinforcement Learning
Yinda Chen, Yangfan He, Jing Yang +5
Prompt engineering significantly influences the reliability and clinical utility of Large Language Models (LLMs) in medical applications. Current optimization approaches inadequate…
SED-MVS: Segmentation-Driven and Edge-Aligned Deformation Multi-View Stereo with Depth Restoration and Occlusion Constraint
Zhenlong Yuan, Zhidong Yang, Yujun Cai +6
Recently, patch-deformation methods have exhibited significant effectiveness in multi-view stereo owing to the deformable and expandable patches in reconstructing textureless areas…
Light4GS: Lightweight Compact 4D Gaussian Splatting Generation via Context Model
Mufan Liu, Qi Yang, He Huang +4
3D Gaussian Splatting (3DGS) has emerged as an efficient and high-fidelity paradigm for novel view synthesis. To adapt 3DGS for dynamic content, deformable 3DGS incorporates tempor…