ChatGPT Images 2.5 on Forgery Tasks: Testing Advertised Improvements Against Known Answers
arXiv:2609.13617
Abstract
OpenAI released ChatGPT Images 2.5 on 8 September 2026, advertising more precise local edits, better consistency across edits, more faithful reference products and sharper detail. We evaluate these claims on four forgery tasks with answers fixed in advance: receipt-field alteration, repeated editing, product placement and small-print rendering. GPT-Image-2 provides same-week baselines at a cheaper and a more expensive tier. A limited improvement appears in receipt editing. After alignment, OCR detects changes to surrounding text in 31.7% of Flare outputs, against 44.2% for the cheaper baseline. This gain is concentrated on CORD receipts and sensitive to shifts of a pixel or less; the forged value itself is no more often correct. Repeated editing and fine print show no measurable gain. Product codes become more legible mainly because Images 2.5 draws the product larger. Defence outcomes change little: localisation remains weak for both generations. A detector that flags 68.6% of controlled benchmark images flags only 35.9% of images posted online. Advertised improvements therefore transfer unevenly to the tested forgery capabilities, while substantial detection limitations remain.
27 pages, 6 figures, 16 tables