◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Qi Xu

9 papers hereh-index 4155 citations9 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author8
  • last author1

Across the 9 of 9 papers where every author was matched, so the position is known.

fields
  • cs.CV4
  • cs.CL3
  • cs.AI2
same name
  • Qi Xu — 8 papers, h 10
  • Qi Xu — 8 papers, h 4
  • Qi Xu — 6 papers, h 12
  • Qi Xu — 6 papers, h 2
  • Qi Xu — 6 papers, h 3
  • Qi Xu — 6 papers, h 4

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedAdvancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

1 citations · 3 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

4 papers · 1 filter

cs.CV2026

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

Jiahao Meng, Yue Tan, Qi Xu +12

Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video…

cs.CV2026

VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification

Jiahao Meng, Tan Yue, Qi Xu +7

Recent video multimodal large language models achieve impressive results across various benchmarks. However, current evaluations suffer from two critical limitations: (1) inflated…

cs.CV2024★ 1 cited

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

Wei Wang, Zhaowei Li, Qi Xu +7

Multi-modal large language models (MLLMs) have achieved remarkable success in fine-grained visual understanding across a range of tasks. However, they often encounter significant c…

cs.CV2024★ 1 cited

GroundingGPT:Language Enhanced Multi-modal Grounding Model

Zhaowei Li, Qi Xu, Dong Zhang +9

Multi-modal large language models have demonstrated impressive performance across various tasks in different modalities. However, existing multi-modal models primarily emphasize ca…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.