◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Haifeng Wang

28 papers hereh-index 6209 citations40 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author6
  • last author22

Across the 28 of 28 papers where every author was matched, so the position is known.

fields
  • cs.CL16
  • cs.CV4
  • cs.SD3
  • cs.AI2
  • cs.LG2
  • cs.SE1
same name
  • Haifeng Wang — 12 papers, h 15
  • Haifeng Wang — 7 papers, h 4
  • Haifeng Wang — 5 papers, h 3
  • Haifeng Wang — 4 papers, h 2
  • Haifeng Wang — 3 papers, h 2
  • Haifeng Wang — 3 papers, h 5

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators
Showing cs.CVShow all

4 papers · 1 filter

cs.CV2026

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

Yuchen Feng, Zhenyu Zhang, Naibin Gu +8

Multimodal large language models (MLLMs) have achieved remarkable progress on various vision-language tasks, yet their visual perception remains limited. Humans, in comparison, per…

cs.CV2026

Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models

Jiadong Pan, Liang Li, Yuxin Peng +6

Recently, unified multimodal models (UMMs) have made remarkable progress in integrating visual understanding and generation, demonstrating strong potential for complex text-to-imag…

cs.CV2026

VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction

Longbin Ji, Xiaoxiong Liu, Junyuan Shang +4

Recent advances in video generation have been dominated by diffusion and flow-matching models, which produce high-quality results but remain computationally intensive and difficult…

cs.CV2025

V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention

Nan Sun, Zhenyu Zhang, Xixun Lin +8

Multimodal Large Language Models (MLLMs) excel in numerous vision-language tasks yet suffer from hallucinations, producing content inconsistent with input visuals, that undermine r…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.