most citedSeePerSea: Multi-modal Perception Dataset of In-water Objects for Autonomous Surface Vehicles

7 citations · 7 across the 5 of their papers we have counts for

collaborators

7 papers

cs.RO2026

Guava: An Effective and Universal Harness for Embodied Manipulation

Haowen Liu, Xirui Li, Shaoxiong Yao +5

Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising…

cs.MA2026

Bridging Perception and Action: A Lightweight Multimodal Meta-Planner Framework for Robust Earth Observation Agents

Jinghui Xu, Boyi Shangguan, Mengke Zhu +10

Autonomous Earth Observation (EO) agents are transitioning from passive perception to complex, multi-step task execution. However, current architectures that integrate planning and…

cs.CV2026

Information Coordination as a Bridge: A Neuro-Symbolic Architecture for Reliable Autonomous Driving Scene Understanding

Shuo Liu, Lei Shi, Haowen Liu +3

Reliable autonomous driving requires scene understanding that is semantically consistent across heterogeneous sensors and verifiable at the reasoning stage. However, many recent LL…

cs.RO2026

SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models

Haowen Liu, Shaoxiong Yao, Haonan Chen +4

Vision-Language Models (VLMs) exhibit remarkable common-sense and semantic reasoning capabilities. However, they lack a grounded understanding of physical dynamics. This limitation…

cs.RO20267 cited

SeePerSea: Multi-modal Perception Dataset of In-water Objects for Autonomous Surface Vehicles

Mingi Jeong, Arihant Chadda, Ziang Ren +8

This paper introduces the first publicly accessible labeled multi-modal perception dataset for autonomous maritime navigation, focusing on in-water obstacles within the aquatic env…

cs.CV2025

PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos

Ting-Hsuan Liao, Haowen Liu, Yiran Xu +3

We present PAD3R, a method for reconstructing deformable 3D objects from casually captured, unposed monocular videos. Unlike existing approaches, PAD3R handles long video sequences…