activity
20242026
collaborators
Showing cs.ROShow all

5 papers · 1 filter

cs.RO2026

See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation

Tingjun Dai, Mingfei Han, Tingwen Du +6

Measurement of task progress through explicit, actionable milestones is critical for robust robotic manipulation. This progress awareness enables a model to ground its current task…

cs.RO2026

Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation

Minghao Jin, Mozheng Liao, Mingfei Han +2

Recent world-model-based Vision-Language-Action (VLA) architectures have improved robotic manipulation through predictive visual foresight. However, dense future prediction introdu…

cs.RO2025

PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly

Liang Ma, Jiajun Wen, Min Lin +12

While vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particul…

cs.RO2025

BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation

Rocktim Jyoti Das, Harsh Singh, Diana Turmakhan +5

Scaling data and models has played a pivotal role in the remarkable progress of computer vision and language. Inspired by these domains, recent efforts in robotics have similarly f…

cs.RO2025

MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation

Harsh Singh, Rocktim Jyoti Das, Mingfei Han +2

Large Language Models (LLMs) have demonstrated remarkable planning abilities across various domains, including robotics manipulation and navigation. While recent efforts in robotic…