8 citations · 19 across the 26 of their papers we have counts for
18 papers · 1 filter
LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation
Yixuan Ding, Jiahao Kong, Wei Huang +2
Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent context. While recency-based caching preserves local continuity, it…
Beyond Pixels: From Video Priors to 4D Worlds
Zihao Liu, Xiaolong Shen, Zhenglin Zhou +2
4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a par…
SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation
Xu Zhang, Yu Lu, Ruijie Quan +3
Flow Matching has enabled robust text-to-video generation via latent ODE sampling. However, velocity approximation and numerical discretization errors inevitably accumulate, causin…
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models
Hong Jiang, Wensong Song, Zongxin Yang +2
Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geometric consistency. However, ex…
Echoes of ownership: Adversarial-guided dual injection for copyright protection in MLLMs
Chengwei Xia, Fan Ma, Ruijie Quan +3
With the rapid deployment of multimodal large language models (MLLMs), disputes regarding model ownership have become increasingly frequent, raising significant concerns about inte…
Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
Xu Zhang, Ruijie Quan, Wenguan Wang +1
Reconstructing visual stimuli from fMRI signals is a central challenge bridging machine learning and neuroscience. Recent diffusion-based methods typically map fMRI activity to a s…