2 papers
cs.CV2026
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
Zhiyuan Ma, Zhengfeng Shi, Yuning An +6
While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semantic alignment and logical rea…
cs.CV2026
I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing
Jinghan Yu, Junhao Xiao, Chenyu Zhu +9
Existing text-guided image editing methods primarily rely on end-to-end pixel-level inpainting paradigm. Despite its success in simple scenarios, this paradigm still significantly…