2 papers
cs.CV2026
NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation
Jiawei Zhang, Shuhao Liu, Rong Huang +4
One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade…
cs.CV2025
Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
Changlin Li, Jiawei Zhang, Shuhao Liu +4
Human video generation has advanced rapidly with the development of diffusion models, but the high computational cost and substantial memory consumption associated with training th…