activity
20192026
most citedWhen language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis

43 citations · 59 across the 11 of their papers we have counts for

collaborators
Showing 2024Show all

6 papers · 1 filter

cs.CV2024

Multiview Scene Graph

Juexiao Zhang, Gao Zhu, Sihang Li +4

A proper scene representation is central to the pursuit of spatial intelligence where agents can robustly reconstruct and efficiently understand 3D scenes. A scene representation i…

cs.CV2024

CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

Xinhao Liu, Jintong Li, Yicheng Jiang +6

Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progres…

cs.RO2024★ 11 cited

VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model

Beichen Wang, Juexiao Zhang, Shuwen Dong +2

Vision Language Models (VLMs) have recently been adopted in robotics for their capability in common sense reasoning and generalizability. Existing work has applied VLMs to generate…

cs.CV2024★ 1 cited

Tell Me Where You Are: Multimodal LLMs Meet Place Recognition

Zonglin Lyu, Juexiao Zhang, Mingxuan Lu +2

Large language models (LLMs) exhibit a variety of promising capabilities in robotics, including long-horizon planning and commonsense reasoning. However, their performance in place…

cs.CV2024★ 4 cited

LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic Images

Jing Zhang, Irving Fang, Juexiao Zhang +7

Lithic Use-Wear Analysis (LUWA) using microscopic images is an underexplored vision-for-science research area. It seeks to distinguish the worked material, which is critical for un…

cs.CV2024

ActFormer: Scalable Collaborative Perception via Active Queries

Suozhi Huang, Juexiao Zhang, Yiming Li +1

Collaborative perception leverages rich visual observations from multiple robots to extend a single robot's perception ability beyond its field of view. Many prior works receive me…