activity
20242026
collaborators

7 papers

cs.CV2026

SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

Yuyang Zhao, Yicheng Pan, Qiyuan He +6

Real-time streaming video-to-video editing (V2V) is critical for interactive applications such as live broadcasting and gaming, yet it remains a formidable challenge due to the str…

cs.CV2026

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation

Xinyao Liao, Qiyuan He, Yicong Li +4

Autoregressive image and video generators are trained with teacher-forced histories but must sample from their own generated prefixes at inference time, making them vulnerable to e…

cs.CV2026

Does Engram Do Memory Retrieval in Autoregressive Image Generation?

Jinghao Wang, Qiyuan He, Chunbin Gu +1

The Engram module -- a hash-keyed, O(1) associative memory injected into Transformer layers -- was recently shown to improve large language model pretraining, with the appealing in…

cs.CV2026

RelaxFlow: Text-Driven Amodal 3D Generation

Jiayin Zhu, Guoji Fu, Xiaolu Liu +3

Image-to-3D generation faces inherent semantic ambiguity under occlusion, where partial observation alone is often insufficient to determine object category. In this work, we forma…

cs.CV2025

CoCA: Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

Xinyao Liao, Wei Wei, Xiaoye Qu +3

Recent advances in text-to-image (T2I) diffusion model fine-tuning leverage reinforcement learning (RL) to align generated images with learnable reward functions. The existing appr…

cs.CV2024

Retinexmamba: Retinex-based Mamba for Low-light Image Enhancement

Jiesong Bai, Yuhao Yin, Qiyuan He +2

In the field of low-light image enhancement, both traditional Retinex methods and advanced deep learning techniques such as Retinexformer have shown distinct advantages and limitat…