3 papers
cs.CV2026
ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations
Junke Wang, Xiao Wang, Jiacheng Pan +16
This paper introduces ARM, a discrete representation-based AutoRegressive Model that unifies image understanding, generation, and editing within a next-token prediction framework.…
cs.CV2025
AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
Arman Zarei, Jiacheng Pan, Matthew Gwilliam +2
Text-to-image generative models have achieved remarkable visual quality but still struggle with compositionalityaccurately capturing object relationships, attribute bindings, an…
cs.CL2025
Improving Text-to-Image Generation with Input-Side Inference-Time Scaling
Ruibo Chen, Jiacheng Pan, Heng Huang +1
Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models often struggle with simple or underspecified prompts, leading to suboptimal…