activity
20242026
collaborators

12 papers

cs.CV2026

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

Junke Wang, Xiao Wang, Jiacheng Pan +16

This paper introduces ARM, a discrete representation-based AutoRegressive Model that unifies image understanding, generation, and editing within a next-token prediction framework.…

cs.CV2026

OmniGen-AR: AutoRegressive Any-to-Image Generation

Junke Wang, Xun Wang, Qiushan Guo +4

Autoregressive (AR) models have demonstrated strong potential in visual generation, offering superior performance with simple architectures and optimization objectives. However, ex…

cs.CV2026

Leveraging Verifier-Based Reinforcement Learning in Image Editing

Hanzhong Guo, Jie Wu, Jie Liu +6

While Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm for text-to-image generation, its application to image editing remains largely unexplored. A k…

cs.CV2025

Seedream 4.0: Toward Next-generation Multimodal Image Generation

Team Seedream, :, Yunpeng Chen +48

We introduce Seedream 4.0, an efficient and high-performance multimodal image generation system that unifies text-to-image (T2I) synthesis, image editing, and multi-image compositi…

cs.LG2025

Scaling Diffusion Transformers Efficiently via P

Chenyu Zheng, Xinyu Zhang, Rongzhen Wang +5

Diffusion Transformers have emerged as the foundation for vision generative models, but their scalability is limited by the high cost of hyperparameter (HP) tuning at large scales.…

cs.AI2025

MEF: A Systematic Evaluation Framework for Text-to-Image Models

Xiaojing Dong, Weilin Huang, Liang Li +6

Rapid advances in text-to-image (T2I) generation have raised higher requirements for evaluation methodologies. Existing benchmarks center on objective capabilities and dimensions,…