Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
ReViCo: Unveiling the Limitations of VLMs in Visual Text Understanding via Error Correction
Bojun Zhang, Junhong Liang, Feifei Zhai +2
Vision Language Models (VLMs) have shown great success in general visual tasks, yet they still struggle to deeply understand text within images. In this paper, we introduce ReViCo…
cs.CV2026
ServImage: An Image Generation and Editing Benchmark from Real-world Commercial Imaging Services
Fengxian Ji, Jingpu Yang, Zirui Song +5
Recent image generation and editing models demonstrate robust adherence to instructions and high visual quality on academic benchmarks. However, their performance on paid, real-wor…