2 papers
cs.CV2026
Attribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models
Xindi Yang, Yicheng Wu, Cheng Zhang +2
Autoregressive text-to-image generation has recently achieved remarkable progress, offering high-fidelity synthesis via a unified generative framework. However, fine-grained semant…
cs.CV2025
VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
Xindi Yang, Baolu Li, Yiming Zhang +8
Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their po…