most citedGRaD-Nav++: Vision-Language Model Enabled Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics

1 citations · 1 across the 6 of their papers we have counts for

collaborators

10 papers

cs.RO2026

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation

Justin Yu, Andrew Goldberg, Kavish Kondap +7

Scaling imitation learning requires large datasets, yet human teleoperation inevitably produces mixed-quality demonstrations containing hesitations and recoveries. Prior frame-leve…

cs.RO2026

Pose-Agnostic Robotic Functional Grasping via Observation-Action Canonicalization

Le Qiu, Cole Harrison, Jiankai Sun +5

Functional robotic grasping requires a policy that generalizes across diverse object geometries and poses while maintaining task-specific contact precision. We study this challenge…

cs.RO2026

SARM2: Multi-Task Stage Aware Reward Modeling for Self Improving Robotic Manipulation

Qianzhong Chen, Hau Zheng, Justin Yu +8

Fine-tuning vision-language-action (VLA) policies for long-horizon manipulation still relies heavily on behavior cloning, which requires costly high-quality demonstrations and keep…

cs.RO20261 cited

GRaD-Nav++: Vision-Language Model Enabled Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics

Qianzhong Chen, Naixiang Gao, Suning Huang +4

Autonomous drones capable of interpreting and executing high-level language instructions in unstructured environments remain a long-standing goal. Yet existing approaches are const…

cs.RO2026

SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation

Qianzhong Chen, Justin Yu, Mac Schwager +3

Large-scale robot learning has made progress on complex manipulation tasks, yet long horizon, contact rich problems, especially those involving deformable objects, remain challengi…

cs.RO2026

Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training

Suning Huang, Jiaqi Shao, Ke Wang +5

Have you ever post-trained a generalist vision-language-action (VLA) policy on a small demonstration dataset, only to find that it stops responding to new instructions and is limit…