◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Chen Gao

20 papers hereh-index 7197 citations26 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author20

Across the 20 of 20 papers where every author was matched, so the position is known.

fields
  • cs.CV9
  • cs.RO7
  • cs.AI4
same name
  • Chen Gao — 14 papers, h 9
  • Chen Gao — 12 papers, h 10
  • Chen Gao — 10 papers, h 3
  • Chen Gao — 9 papers, h 3
  • Chen Gao — 9 papers, h 4
  • Chen Gao — 6 papers, h 3

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.AIShow all

4 papers · 1 filter

cs.AI2026

WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation

Shengtao Zheng, Kai Li, Weichen Zhang +5

End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation. However, existing approaches typically rely on historical observations to directly predict acti…

cs.AI2026

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace

Baining Zhao, Ziyou Wang, Jianjie Fang +8

Large multimodal models (LMMs) show strong visual-linguistic reasoning but their capacity for spatial decision-making and action remains unclear. In this work, we investigate wheth…

cs.AI2026

WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models

Hongjin Chen, Shangyun Jiang, Tonghua Su +4

Vision-language models (VLMs) and generative world models are opening new opportunities for embodied navigation. VLMs are increasingly used as direct planners or trajectory predict…

cs.AI2025

Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning

Baining Zhao, Ziyou Wang, Jianjie Fang +7

Humans can perceive and reason about spatial relationships from sequential visual observations, such as egocentric video streams. However, how pretrained models acquire such abilit…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.