◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Qifan Wang

17 papers hereh-index 9347 citations20 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author17

Across the 17 of 17 papers where every author was matched, so the position is known.

fields
  • cs.CL9
  • cs.CV4
  • cs.AI3
  • cs.MA1
same name
  • Qifan Wang — 22 papers, h 18
  • Qifan Wang — 17 papers, h 9
  • Qifan Wang — 14 papers, h 8
  • Qifan Wang — 10 papers, h 7
  • Qifan Wang — 10 papers, h 5
  • Qifan Wang — 9 papers, h 6

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20232026
most citedVision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

2 citations · 4 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

4 papers · 1 filter

cs.CV2025

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

Zhiyang Xu, Jiuhai Chen, Zhaojiang Lin +10

Recent advances in large language models (LLMs) have enabled multimodal foundation models to tackle both image understanding and generation within a unified framework. Despite thes…

cs.CV2025

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation

Jingyuan Qi, Zhiyang Xu, Qifan Wang +1

We introduce Autoregressive Retrieval Augmentation (AR-RAG), a novel paradigm that enhances image generation by autoregressively incorporating knearest neighbor retrievals at the p…

cs.CV2024★ 1 cited

Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Haibo Wang, Zhiyang Xu, Yu Cheng +6

Video Large Language Models (Video-LLMs) have demonstrated remarkable capabilities in coarse-grained video understanding, however, they struggle with fine-grained temporal groundin…

cs.CV2024

Multimodal Instruction Tuning with Conditional Mixture of LoRA

Ying Shen, Zhiyang Xu, Qifan Wang +3

Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in diverse tasks across different domains, with an increasing focus on improving their zero-shot g…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.