◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Hang Li

4 papers hereh-index 375 citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author2
  • last author1

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CV3
  • cs.AI1
same name
  • Hang Li — 18 papers, h 18
  • Hang Li — 14 papers, h 55
  • Hang Li — 9 papers, h 7
  • Hang Li — 9 papers, h 5
  • Hang Li — 8 papers
  • Hang Li — 8 papers, h 17

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedSeeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

1 citations · 1 across the 4 of their papers we have counts for

collaborators

4 papers

cs.CV2026

Geometry-Guided 3D Visual Token Pruning for Video-Language Models

Han Li, Zehao Huang, Jiahui Fu +2

Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent studies represent 3D scenes as…

cs.CV2026

PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues

Yukun Qi, Pei Fu, Hang Li +5

Vision-Language Models (VLMs) have achieved remarkable progress on a wide range of challenging multimodal understanding and reasoning tasks. However, existing reasoning paradigms,…

cs.AI2025

Robix: A Unified Model for Robot Interaction, Reasoning and Planning

Huang Fang, Mengxi Zhang, Heng Dong +6

We introduce Robix, a unified model that integrates robot reasoning, task planning, and natural language interaction within a single vision-language architecture. Acting as the hig…

cs.CV2025★ 1 cited

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

Lin Long, Yichen He, Wentao Ye +5

We introduce M3-Agent, a novel multimodal agent framework equipped with long-term memory. Like humans, M3-Agent can process real-time visual and auditory inputs to build and update…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.