activity
20232026
most citedFusionSAM: Visual Multi-Modal Learning with Segment Anything

3 citations · 5 across the 22 of their papers we have counts for

collaborators
Showing cs.ROShow all

7 papers · 1 filter

cs.RO2026

OccamView: Object-Conditioned View Selection for Frame-Budgeted Active 3D Gaussian Reconstruction

Hongbo Gao, Wei Zhang, Zeyu Ni +4

Active 3D Gaussian reconstruction fundamentally relies on selecting informative next-best views under limited sensing budgets. Existing active 3DGS methods primarily plan viewpoint…

cs.RO2026

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models

Siyu Xu, Yunke Wang, Zijian Wang +6

Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or o…

cs.RO2026

See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model

Yixu Feng, Zinan Zhao, Yanxiang Ma +4

Vision-Language-Action (VLA) models have shown remarkable promise in robotics manipulation, yet their high computational cost hinders real-time deployment. Existing token pruning m…

cs.RO2026

SmoothTurn: Learning to Turn Smoothly for Agile Navigation with Quadrupedal Robots

Zunzhi You, Yunke Wang, Haolan Guo +1

Quadrupedal robots show great potential for valuable real-world applications such as fire rescue and industrial inspection. Such applications often require urgency and the ability…

cs.RO2025

Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation

Siyu Xu, Zijian Wang, Yunke Wang +3

Vision-Language-Action (VLA) models have shown great performance in robotic manipulation by mapping visual observations and language instructions directly to actions. However, they…

cs.RO2025

Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation

Xiaohuan Pei, Yuxing Chen, Siyu Xu +3

Robotic manipulation with Vision-Language-Action models requires efficient inference over long-horizon multi-modal context, where attention to dense visual tokens dominates computa…