6 papers
Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models
Haibo Wang, Lifu Huang
Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations that fail to maintain geometric and spatia…
Open-Ended Video Game Glitch Detection with Agentic Reasoning and Temporal Grounding
Muyang Zheng, Tong Zhou, Geyang Wu +3
Open-ended video game glitch detection aims to identify glitches in gameplay videos, describe them in natural language, and localize when they occur. Unlike conventional game glitc…
GLANCE: A Global-Local Coordination Multi-Agent Framework for Music-Grounded Non-Linear Video Editing
Zihao Lin, Haibo Wang, Zhiyang Xu +7
Music-grounded mashup video creation is a challenging form of video non-linear editing, where a system must compose a coherent timeline from large collections of source videos whil…
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
Haibo Wang, Zihao Lin, Zhiyang Xu +1
3D Visual Grounding (3D-VG) aims to localize objects in 3D scenes via natural language descriptions. While recent advancements leveraging Vision-Language Models (VLMs) have explore…
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
Zhiyang Xu, Tian Qin, Bowen Jin +4
Multimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings…
Zero-shot adaptable task planning for autonomous construction robots: a comparative study of lightweight single and multi-AI agent systems
Hossein Naderi, Alireza Shojaei, Lifu Huang +3
Robots are expected to play a major role in the future construction industry but face challenges due to high costs and difficulty adapting to dynamic tasks. This study explores the…