3 papers
cs.CV2026
MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing
Liwei Cheng, Shibo Feng, Lunjie Zhou +2
Text-in-image editing has become a key capability for visual content creation, yet existing benchmarks remain overwhelmingly English-centric and often conflate visual plausibility…
cs.CV2025
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
Xuan Yu, Dayan Guan, Yanfeng Gu
Multimodal Large Language Models (MLLM) often struggle to interpret high-resolution images accurately, where fine-grained details are crucial for complex visual understanding. We i…
cs.CV2024
Rethinking the Evaluation of Visible and Infrared Image Fusion
Dayan Guan, Yixuan Wu, Tianzhu Liu +2
Visible and Infrared Image Fusion (VIF) has garnered significant interest across a wide range of high-level vision tasks, such as object detection and semantic segmentation. Howeve…