3 papers
cs.CV2026
COLLAR: Cascaded Object-Level Latent Refinement for High-Fidelity Conditional Generation
Xinlong Zhang, Jia Wei, Xiaoyu Zhang +3
Achieving high-fidelity object-level control in Diffusion Transformers remains a significant challenge despite the introduction of structural priors like depth and Canny maps. Curr…
cs.CV2025
PanoLlama: Generating Endless and Coherent Panoramas with Next-Token-Prediction LLMs
Teng Zhou, Xiaoyu Zhang, Yongchuan Tang
Panoramic Image Generation (PIG) aims to create coherent images of arbitrary lengths. Most existing methods fall in the joint diffusion paradigm, but their complex and heuristic cr…
cs.CV2025
Multi-Scale Diffusion: Enhancing Spatial Layout in High-Resolution Panoramic Image Generation
Xiaoyu Zhang, Teng Zhou, Xinlong Zhang +2
Diffusion models have recently gained recognition for generating diverse and high-quality content, especially in image synthesis. These models excel not only in creating fixed-size…