collaborators

7 papers

cs.RO2026

Bridge-WA: Predicting Where and How the World Changes for Robotic Action

Yongjie Bai, Hanting Wang, Mingtong Dai +3

General-purpose vision-language-action models benefit from large vision-language priors, but effective manipulation also requires anticipating action-relevant scene changes. Existi…

cs.LG2026

Fitting Unknown Number of Hyperplanes with Manifold Optimization

Zhiqin Cheng, Yu Zhan, Mingjin Zhang +2

Fitting an unknown number of hyperplanes to data is a fundamental yet challenging problem in machine learning, characterized by its non-convexity, non-differentiability, and unknow…

cs.RO2026

SkiP: When to Skip and When to Refine for Efficient Robot Manipulation

Mingtong Dai, Guanqi Peng, Yongjie Bai +5

Previous imitation learning policies predict future actions at every control step, whether in smooth motion phases or precise, contact-rich operation phases. This uniform treatment…

cs.RO2026

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

Yongjie Bai, Zhouxia Wang, Yang Liu +8

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlu…

cs.CV2026

PointCoT: A Multi-modal Benchmark for Explicit 3D Geometric Reasoning

Dongxu Zhang, Yiding Sun, Pengcheng Li +12

While Multimodal Large Language Models (MLLMs) demonstrate proficiency in 2D scenes, extending their perceptual intelligence to 3D point cloud understanding remains a significant c…

cs.RO2025

GraspView: Active Perception Scoring and Best-View Optimization for Robotic Grasping in Cluttered Environments

Shenglin Wang, Mingtong Dai, Jingxuan Su +4

Robotic grasping is a fundamental capability for autonomous manipulation, yet remains highly challenging in cluttered environments where occlusion, poor perception quality, and inc…