4 papers
Training-Inference Consistent Segmented Execution for Long-Context LLMs
Xianpeng Shang, Jiang Li, Zehua Duo +2
Transformer-based large language models face severe scalability challenges in long-context generation due to the computational and memory costs of full-context attention. Under pra…
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
Qi Cai, Jingwen Chen, Chengmin Gao +22
The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In this report, we present HiDr…
Visual Autoregressive Modeling for Instruction-Guided Image Editing
Qingyang Mao, Qi Cai, Yehao Li +5
Recent advances in diffusion models have brought remarkable visual fidelity to instruction-guided image editing. However, their global denoising process inherently entangles the ed…
Creatively Upscaling Images with Global-Regional Priors
Yurui Qian, Qi Cai, Yingwei Pan +2
Contemporary diffusion models show remarkable capability in text-to-image generation, while still being limited to restricted resolutions (e.g., 1,024 X 1,024). Recent advances ena…