activity
20242026
collaborators
Showing cs.CVShow all

26 papers · 1 filter

cs.CV2026

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

Zitong Xu, Huiyu Duan, Xinyun Zhang +7

Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstr…

cs.CV2026

Towards Photorealistic and Efficient Bokeh Rendering via Diffusion Framework

Linxiao Shi, Siming Zheng, Zerong Wang +5

Existing mobile devices are constrained by compact optical designs, such as small apertures, which make it difficult to produce natural, optically realistic bokeh effects. Although…

cs.CV2026

How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings

Zhiheng Li, Zongyang Ma, Jiaxian Chen +12

The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanually annotated dataset whose t…

cs.CV2026

EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

Zitong Xu, Huiyu Duan, Yifei Nie +9

Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as unnatural objects, lighting…

cs.CV2026

Visual Text Compression as Measure Transport

Lv Tang, Tianyi Zheng, Yang Liu +2

Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language model, often producing --$20\t…

cs.CV2026

HyperMotionX: The Dataset and Benchmark with DiT-Based Pose-Guided Human Image Animation of Complex Motions

Shuolin Xu, Siming Zheng, Ziyi Wang +7

Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods…