8 papers
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
Yiren Song, Huilin Zhong, Kevin Qinghong Lin +2
We study series-level cinematic remaking, a long-horizon video-to-video generation problem that localizes full episodes or films via stylization or actor replacement while strictly…
VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
Yiren Song, Wangzi Yao, Haofan Wang +1
Video style transfer aims to render videos in a target artistic style while preserving content, structure, and motion. While image stylization has advanced rapidly, video stylizati…
Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs
Yiren Zheng, Shibo Li, Jiaming Liu +2
Current multimodal approaches predominantly treat visual generation as an external process, relying on pixel rendering or code execution, thereby overlooking the native visual repr…
SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens
Xiaoyan Zhang, Zechen Bai, Haofan Wang +1
Recent unified models such as Bagel demonstrate that paired image-edit data can effectively align multiple visual tasks within a single diffusion transformer. However, these models…
OmniPSD: Layered PSD Generation with Diffusion Transformer
Cheng Liu, Yiren Song, Haofan Wang +1
Recent advances in diffusion models have greatly improved image generation and editing, yet generating or reconstructing layered PSD files with transparent alpha channels remains h…
GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
Chun Wang, Xiaojun Ye, Xiaoran Pan +3
Recent advances in Visual Language Models (VLMs) have demonstrated exceptional performance in visual reasoning tasks. However, geo-localization presents unique challenges, requirin…