collaborators

5 papers

cs.CV2026

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?

Han Li, Si Liu, Zehao Huang +6

Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturall…

cs.CV2026

Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior

Jiahui Fu, Zehao Huang, Han Li +2

Lane topology reasoning aims to construct a lane graph from onboard sensor observations. Existing methods follow a detection and association paradigm that treats each lane instance…

cs.CV2026

Geometry-Guided 3D Visual Token Pruning for Video-Language Models

Han Li, Zehao Huang, Jiahui Fu +2

Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent studies represent 3D scenes as…

cs.RO2026

NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning

Jiahui Fu, Junyu Nan, Lingfeng Sun +5

Solving long-horizon tasks requires robots to integrate high-level semantic reasoning with low-level physical interaction. While vision-language models (VLMs) and video generation…

cs.RO2025

NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos

Hongyu Li, Lingfeng Sun, Yafei Hu +4

Enabling robots to execute novel manipulation tasks zero-shot is a central goal in robotics. Most existing methods assume in-distribution tasks or rely on fine-tuning with embodime…