3 papers
cs.CV2025
PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models
Wanru Zhuang, Wenbo Li, Zhibin Lan +3
Text Image Machine Translation (TIMT) aims to translate texts embedded within an image into another language. Current TIMT studies primarily focus on providing translations for all…
cs.CV2024
Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration
Haoze Sun, Wenbo Li, Jiayue Liu +7
Generalization has long been a central challenge in real-world image restoration. While recent diffusion-based restoration methods, which leverage generative priors from text-to-im…
cs.CV2024
RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models
Haoyu Chen, Wenbo Li, Jinjin Gu +7
Natural images captured by mobile devices often suffer from multiple types of degradation, such as noise, blur, and low light. Traditional image restoration methods require manual…