Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Consistent text-to-image generation via scene de-contextualization
Song Tang, Peihao Gong, Kunyu Li +5
Consistent text-to-image (T2I) generation seeks to produce identity-preserving images of the same subject across diverse scenes, yet it often fails due to a phenomenon called ident…
cs.CV2025
Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
Jiahuan Zhang, Shunwen Bai, Tianheng Wang +4
Humans naturally possess the spatial reasoning ability to form and manipulate images and structures of objects in space. There is an increasing effort to endow Vision-Language Mode…