activity
20212026
most citedMulti-direction and Multi-scale Pyramid in Transformer for Video-based Pedestrian Retrieval

99 citations · 203 across the 14 of their papers we have counts for

collaborators

15 papers

cs.CV2026

VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal

Haonan Huang, Tianrui Qiu, Xianghao Zang +10

Despite its crucial role in video object removal (VOR), existing evaluation paradigms face two critical limitations: questionable references and a misalignment between tradi- tiona…

cs.CV2026

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

Xianghao Zang, Zijian Jiang, Jiarong Cheng +8

Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Mod…

cs.CV2026

Style-CCL: Content-Preserving Style Transfer via Curriculum Continual Learning

Shiwen Zhang, Haoyuan Wang, Xianghao Zang +3

Content-Preserving Style transfer, given content and style references, remains challenging for Diffusion Transformers (DiTs) due to entangled content and style features. With a rev…

cs.CV2026

Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation

Baoteng Li, Xianghao Zang, Xinran Wang +8

Text-to-Image (T2I) generation has achieved remarkable progress in recent years. Meanwhile, reinforcement learning methods, particularly those based on Group Relative Policy Optimi…

cs.AI2026

DataEvolver: Let Your Data Build and Improve Itself via Goal-Driven Loop Agents

Qisong Zhang, Wenzhuo Wu, Zhuangzhuang Jia +7

Constructing controllable visual data is a major bottleneck for image editing and multimodal understanding. Useful supervision is rarely produced by a single rendering pass; instea…

cs.CV2026

Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion Transformers

Shuo Zhang, Wenzhuo Wu, Huayu Zhang +8

Recent advances in diffusion models have significantly improved image editing. However, challenges persist in handling geometric transformations, such as translation, rotation, and…