◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Honghao Cai

8 papers hereh-index 215 citations11 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author6

Across the 8 of 8 papers where every author was matched, so the position is known.

fields
  • cs.CV4
  • cs.RO4

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.CVShow all

4 papers · 1 filter

cs.CV2026

EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis

Xiangyuan Wang, Honghao Cai, Yunhao Bai +9

High-quality source-target image pairs with precise editing instructions are essential for instruction-guided image editing, yet constructing such training triplets at scale remain…

cs.CV2026

IdGlow: Dynamic Identity Modulation for Multi-Subject Generation

Honghao Cai, Xiangyuan Wang, Jing Li +15

Multi-subject image generation requires seamlessly harmonizing multiple reference identities within a coherent scene. However, existing methods relying on rigid spatial masks or lo…

cs.CV2026

Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing

Honghao Cai, Xiangyuan Wang, Yunhao Bai +8

Large diffusion transformers (DiTs) follow global editing instructions well but consistently leak local edits into unrelated regions, because joint-attention architectures offer no…

cs.CV2026

SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning

Yuyuan Yang, Junkun Hong, Hongrong Wang +11

Embodied task planning demands vision-language models to generate action sequences that are both visually grounded and causally coherent over time. However, existing training parad…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.