collaborators

7 papers

cs.CV2025

2K-Characters-10K-Stories: A Quality-Gated Stylized Narrative Dataset with Disentangled Control and Sequence Consistency

Xingxi Yin, Yicheng Li, Gong Yan +5

Sequential identity consistency under precise transient attribute control remains a long-standing challenge in controllable visual storytelling. Existing datasets lack sufficient f…

cs.CV2025

VideoPro: Adaptive Program Reasoning for Long Video Understanding

Chenglin Li, Feng Han, Yikun Wang +9

Large language models (LLMs) have shown promise in generating program workflows for visual tasks. However, previous approaches often rely on closed-source models, lack systematic r…

cs.CV2025

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding

Chenglin Li, Qianglong Chen, fengtao +1

Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding tasks. However, they continue to struggle with long-form videos because of an ineffici…

cs.CV2025

InstructAttribute: Fine-grained Object Attributes editing with Instruction

Xingxi Yin, Jingfeng Zhang, Yue Deng +3

Text-to-image (T2I) diffusion models are widely used in image editing due to their powerful generative capabilities. However, achieving fine-grained control over specific object at…

cs.CV2025

RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers

Yan Gong, Yiren Song, Yicheng Li +2

Inspired by the in-context learning mechanism of large language models (LLMs), a new paradigm of generalizable visual prompt-based image editing is emerging. Existing single-refere…

cs.CV2024

ColorEdit: Training-free Image-Guided Color editing with diffusion model

Xingxi Yin, Zhi Li, Jingfeng Zhang +2

Text-to-image (T2I) diffusion models, with their impressive generative capabilities, have been adopted for image editing tasks, demonstrating remarkable efficacy. However, due to a…