From the 1 of 8 linked papers with an AI index.
8 papers
DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
Jiaxing Li, Kai Zou, Cindy Zhou +7
The paper studies autoregressive video distillation, showing that aligning the student model’s mode coverage with the teacher’s distribution improves generation quality and diversi…
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
Zile Wang, Zexiang Liu, Jiaxing Li +20
With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing approaches still struggle t…
Advances in GRPO for Generation Models: A Survey
Zexiang Liu, Xianglong He, Yangguang Li
Large-scale flow matching models have achieved strong performance across generative tasks such as text-to-image, video, 3D, and speech synthesis. However, aligning their outputs wi…
Faithful Contouring: Near-Lossless 3D Voxel Representation Free from Iso-surface
Yihao Luo, Xianglong He, Chuanyu Pan +7
Accurate and efficient voxelized representations of 3D meshes are the foundation of 3D reconstruction and generation. However, existing representations based on iso-surface heavily…
Transition Models: Rethinking the Generative Learning Objective
Zidong Wang, Yiyuan Zhang, Xiaoyu Yue +4
A fundamental dilemma in generative modeling persists: iterative diffusion models achieve outstanding fidelity, but at a significant computational cost, while efficient few-step al…
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction
Xuying Zhang, Yutong Liu, Yangguang Li +8
We present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQ-VAE) and a Generative Pre-trained Transformer (GPT) to generate high-qu…