4 papers · 1 filter
TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
Rui Gui, Yang Wan, Haochen Han +4
Text rendering has recently emerged as one of the most challenging frontiers in visual generation, drawing significant attention from large-scale diffusion and multimodal models. H…
Negation-Aware Test-Time Adaptation for Vision-Language Models
Haochen Han, Alex Jinpeng Wang, Fangming Liu +1
In this paper, we study a practical but less-touched problem in Vision-Language Models (VLMs), \ie, negation understanding. Specifically, many real-world applications require model…
Unlearning the Noisy Correspondence Makes CLIP More Robust
Haochen Han, Alex Jinpeng Wang, Peijun Ye +1
The data appetite for Vision-Language Models (VLMs) has continuously scaled up from the early millions to billions today, which faces an untenable trade-off with data quality and i…
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
Zhuobai Dong, Junchao Yi, Ziyuan Zheng +5
Understanding the physical world - governed by laws of motion, spatial relations, and causality - poses a fundamental challenge for multimodal large language models (MLLMs). While…