collaborators

14 papers

cs.CV2026

EditTransfer++: Toward Faithful and Efficient Visual-Prompt-Guided Image Editing

Lan Chen, Qi Mao, Yiren Song +2

Visual-prompt-guided edit transfer aims to learn image transformations directly from example pairs, offering more precise and controllable editing than purely text-driven approache…

cs.CV2025

Mitty: Diffusion-based Human-to-Robot Video Generation

Yiren Song, Cheng Liu, Weijia Mao +1

Learning directly from human demonstration videos is a key milestone toward scalable and generalizable robot learning. Yet existing methods rely on intermediate representations suc…

cs.CV2025

IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning

Yuanhang Li, Yiren Song, Junzhe Bai +4

We propose \textbf{IC-Effect}, an instruction-guided, DiT-based framework for few-shot video VFX editing that synthesizes complex effects (\eg flames, particles and cartoon charact…

cs.RO2025

H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos

Hai Ci, Xiaokang Liu, Pei Yang +2

Robots that learn manipulation skills from everyday human videos could acquire broad capabilities without tedious robot data collection. We propose a video-to-video translation fra…

cs.CV2025

OmniPSD: Layered PSD Generation with Diffusion Transformer

Cheng Liu, Yiren Song, Haofan Wang +1

Recent advances in diffusion models have greatly improved image generation and editing, yet generating or reconstructing layered PSD files with transparent alpha channels remains h…

cs.CV2025

The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment

Ziheng Ouyang, Yiren Song, Yaoli Liu +4

Previous works have explored various customized generation tasks given a reference image, but they still face limitations in generating consistent fine-grained details. In this pap…