1 citations · 1 across the 4 of their papers we have counts for
6 papers
Pose-Agnostic Robotic Functional Grasping via Observation-Action Canonicalization
Le Qiu, Cole Harrison, Jiankai Sun +5
Functional robotic grasping requires a policy that generalizes across diverse object geometries and poses while maintaining task-specific contact precision. We study this challenge…
SARM2: Multi-Task Stage Aware Reward Modeling for Self Improving Robotic Manipulation
Qianzhong Chen, Hau Zheng, Justin Yu +8
Fine-tuning vision-language-action (VLA) policies for long-horizon manipulation still relies heavily on behavior cloning, which requires costly high-quality demonstrations and keep…
GRaD-Nav++: Vision-Language Model Enabled Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics
Qianzhong Chen, Naixiang Gao, Suning Huang +4
Autonomous drones capable of interpreting and executing high-level language instructions in unstructured environments remain a long-standing goal. Yet existing approaches are const…
Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training
Suning Huang, Jiaqi Shao, Ke Wang +5
Have you ever post-trained a generalist vision-language-action (VLA) policy on a small demonstration dataset, only to find that it stops responding to new instructions and is limit…
Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks
Jindong Hong, Tianjie Chen, Lingjie Luo +10
A recent advancement in Multimodal Large Language Models (MLLMs) research is the emergence of "reasoning MLLMs" that offer explicit control over their internal thinking processes (…
ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
Suning Huang, Qianzhong Chen, Xiaohan Zhang +2
3D world models (i.e., learning-based 3D dynamics models) offer a promising approach to generalizable robotic manipulation by capturing the underlying physics of environment evolut…