6 citations · 7 across the 20 of their papers we have counts for
33 papers
Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
Hengyuan Xu, Wei Cheng, Yumeng Ji +4
Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an imag…
Motion4Motion: Motion Transfer Across Subjects at Inference
Ling-Hao Chen, Zixin Yin, Duomin Wang +2
This work explores the motion transfer from one video to another, which is crucial in animation for diverse characters. Previously, video motion transfer has been largely explored…
FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining
Jinghong Lan, Wei Cheng, Yunuo Chen +10
Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while adopting the style of a separate style r…
ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement
Yufeng Yang, Jianzhuang Liu, Jisheng Chu +4
Existing deep learning-based low-light enhancement methods are typically trained on limited datasets with single enhancement targets, which restricts their generalization ability a…
PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset
Haojun Chen, Haoyang He, Chengming Xu +11
Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the rapid development of imagin…
GEditBench v2: A Human-Aligned Benchmark for General Image Editing
Zhangqi Jiang, Zheng Sun, Xianfang Zeng +7
Recent advances in image editing have enabled models to handle complex instructions with impressive realism. However, existing evaluation frameworks lag behind: current benchmarks…