1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CV2026
Geometry-Guided 3D Visual Token Pruning for Video-Language Models
Han Li, Zehao Huang, Jiahui Fu +2
Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent studies represent 3D scenes as…
cs.RO2026
NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning
Jiahui Fu, Junyu Nan, Lingfeng Sun +5
Solving long-horizon tasks requires robots to integrate high-level semantic reasoning with low-level physical interaction. While vision-language models (VLMs) and video generation…
cs.RO2025★ 1 cited
NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos
Hongyu Li, Lingfeng Sun, Yafei Hu +4
Enabling robots to execute novel manipulation tasks zero-shot is a central goal in robotics. Most existing methods assume in-distribution tasks or rely on fine-tuning with embodime…