49 citations · 53 across the 15 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
Lijie Fan, Luming Tang, Siyang Qin +11
We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture p…
cs.CV2025★ 1 cited
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Guoqing Ma, Haoyang Huang, Kun Yan +112
We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression…
cs.CV2025★ 3 cited
Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
Nanye Ma, Shangyuan Tong, Haolin Jia +8
Generative models have made significant impacts across various domains, largely due to their ability to scale during training by increasing data, computational resources, and model…