◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Yan Zhang

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author1

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CV4
same name
  • Yan Zhang — 54 papers
  • Yan Zhang — 38 papers, h 76
  • Yan Zhang — 23 papers
  • Yan Zhang — 15 papers
  • Yan Zhang — 15 papers
  • Yan Zhang — 13 papers, h 61

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedTrack the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues

1 citations · 1 across the 2 of their papers we have counts for

collaborators

4 papers

cs.CV2025

Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective

Yan Zhang, Gangyan Zeng, Daiqing Wu +5

Video text-based visual question answering (Video TextVQA) aims to answer questions by explicitly reading and reasoning about the text involved in a video. Most works in this field…

cs.CV2025

When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

Yan Shu, Hangui Lin, Yexin Liu +7

Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, th…

cs.CV2025

VidText: Towards Comprehensive Evaluation for Video Text Understanding

Zhoufaran Yang, Yan Shu, Jing Wang +8

Visual texts embedded in videos carry rich semantic information, which is crucial for both holistic video understanding and fine-grained reasoning about local human actions. Howeve…

cs.CV2024★ 1 cited

Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues

Yan Zhang, Gangyan Zeng, Huawen Shen +3

Video text-based visual question answering (Video TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. I…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.