collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video

Weili Guan, Haoyu Zhang, Meng Liu +3

Visual-spatial understanding, defined as the ability to infer object relationships and scene layouts from visual inputs, is fundamental to downstream tasks such as robotic navigati…

cs.CV2026

SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion

Haoyu Zhang, Meng Liu, Qianlong Xiang +3

Spatial intelligence is essential for low-altitude unmanned aerial vehicle (UAV) perception, collaboration, and navigation. However, existing UAV benchmarks often emphasize image-l…

cs.CV2025

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

Haoyu Zhang, Qiaohui Chu, Meng Liu +3

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Mode…

cs.CV2025

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data

Haoyu Zhang, Meng Liu, Zaijing Li +4

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied in…

cs.CV2025

Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025

Qiaohui Chu, Haoyu Zhang, Yisen Feng +4

In this report, we present a novel three-stage framework developed for the Ego4D Long-Term Action Anticipation (LTA) task. Inspired by recent advances in foundation models, our met…

cs.CV2025

OSGNet @ Ego4D Episodic Memory Challenge 2025

Yisen Feng, Haoyu Zhang, Qiaohui Chu +4

In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise…