10 papers
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm
Yaofang Liu, Kangning Cui, Meng Chu +7
Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual generators still ask users to seria…
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
Yaofang Liu, Yumeng Ren, Aitor Artola +9
The rapid advancement of video diffusion models has been hindered by fundamental limitations in temporal modeling, particularly the rigid synchronization of frame evolution imposed…
Seg-Agent: Test-Time Multimodal Reasoning for Training-Free Language-Guided Segmentation
Chao Hao, Jun Xu, Ji Du +6
Language-guided segmentation transcends the scope limitations of traditional semantic segmentation, enabling models to segment arbitrary target regions based on natural language in…
CutClaw: Agentic Hours-Long Video Editing via Music Synchronization
Shifang Zhao, Yihan Hu, Ying Shan +2
Editing the video content with audio alignment forms a digital human-made art in current social media. However, the time-consuming and repetitive nature of manual video editing has…
LightCtrl: Training-free Controllable Video Relighting
Yizuo Peng, Xuelin Chen, Kai Zhang +1
Recent diffusion models have achieved remarkable success in image relighting, and this success has quickly been extended to video relighting. However, existing methods offer limite…
MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence
Xingyilang Yin, Chengzhengxu Li, Jiahao Chang +2
Humans are born with vision-based 4D spatial-temporal intelligence, which enables us to perceive and reason about the evolution of 3D space over time from purely visual inputs. Des…