7 citations · 7 across the 5 of their papers we have counts for
7 papers
Guava: An Effective and Universal Harness for Embodied Manipulation
Haowen Liu, Xirui Li, Shaoxiong Yao +5
Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising…
Bridging Perception and Action: A Lightweight Multimodal Meta-Planner Framework for Robust Earth Observation Agents
Jinghui Xu, Boyi Shangguan, Mengke Zhu +10
Autonomous Earth Observation (EO) agents are transitioning from passive perception to complex, multi-step task execution. However, current architectures that integrate planning and…
Information Coordination as a Bridge: A Neuro-Symbolic Architecture for Reliable Autonomous Driving Scene Understanding
Shuo Liu, Lei Shi, Haowen Liu +3
Reliable autonomous driving requires scene understanding that is semantically consistent across heterogeneous sensors and verifiable at the reasoning stage. However, many recent LL…
SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
Haowen Liu, Shaoxiong Yao, Haonan Chen +4
Vision-Language Models (VLMs) exhibit remarkable common-sense and semantic reasoning capabilities. However, they lack a grounded understanding of physical dynamics. This limitation…
SeePerSea: Multi-modal Perception Dataset of In-water Objects for Autonomous Surface Vehicles
Mingi Jeong, Arihant Chadda, Ziang Ren +8
This paper introduces the first publicly accessible labeled multi-modal perception dataset for autonomous maritime navigation, focusing on in-water obstacles within the aquatic env…
PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos
Ting-Hsuan Liao, Haowen Liu, Yiran Xu +3
We present PAD3R, a method for reconstructing deformable 3D objects from casually captured, unposed monocular videos. Unlike existing approaches, PAD3R handles long video sequences…