1 paper
Yu Huo, Siyu Zhang, Kun Zeng +7
Multimodal models for text-to-image generation have achieved strong visual fidelity, yet they remain brittle under compositional structural constraints, notably generative numeracy…