2 papers
cs.CV2025
Plan-X: Instruct Video Generation via Semantic Planning
Lun Huang, You Xie, Hongyi Xu +7
Diffusion Transformers have demonstrated remarkable capabilities in visual synthesis, yet they often struggle with high-level semantic reasoning and long-horizon planning. This lim…
cs.CV2025
Autoregressive Models in Vision: A Survey
Jing Xiong, Gongye Liu, Lun Huang +17
Autoregressive modeling has been a huge success in the field of natural language processing (NLP). Recently, autoregressive models have emerged as a significant area of focus in co…