5 citations · 5 across the 3 of their papers we have counts for
1 paper · 1 filter
Ashritha Gonuguntla
Reasoning-augmented text-to-image models such as GoT-R1 emit an explicit textual plan - object names, attributes, and bounding boxes - before generating image tokens. When such a m…