6 papers · 1 filter
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
Han Li, Si Liu, Zehao Huang +6
Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturall…
Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior
Jiahui Fu, Zehao Huang, Han Li +2
Lane topology reasoning aims to construct a lane graph from onboard sensor observations. Existing methods follow a detection and association paradigm that treats each lane instance…
VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation
Yulu Gao, Bohao Zhang, Zongheng Tang +3
Instance-level object segmentation across disparate egocentric and exocentric views is a fundamental challenge in visual understanding, critical for applications in embodied AI and…
Geometry-Guided 3D Visual Token Pruning for Video-Language Models
Han Li, Zehao Huang, Jiahui Fu +2
Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent studies represent 3D scenes as…
LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation
Yuwei Ning, Ganlong Zhao, Yipeng Qin +4
Aerial Vision-and-Language Navigation (Aerial VLN) enables unmanned aerial vehicles (UAVs) to follow natural language instructions and navigate complex urban environments. While re…
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
Team HY-World, Chenjie Cao, Xuhui Zuo +42
We introduce HY-World 2.0, a multi-modal world model framework that advances our prior project HY-World 1.0. HY-World 2.0 accommodates diverse input modalities, including text prom…