4 papers
TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
Rui Gui, Yang Wan, Haochen Han +4
Text rendering has recently emerged as one of the most challenging frontiers in visual generation, drawing significant attention from large-scale diffusion and multimodal models. H…
Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification
Rifen Lin, Alex Jinpeng Wang, Jiawei Mo +1
Multimodal pretraining has revolutionized visual understanding, but its impact on video-based person re-identification (ReID) remains underexplored. Existing approaches often rely…
Negation-Aware Test-Time Adaptation for Vision-Language Models
Haochen Han, Alex Jinpeng Wang, Fangming Liu +1
In this paper, we study a practical but less-touched problem in Vision-Language Models (VLMs), \ie, negation understanding. Specifically, many real-world applications require model…
Unlearning the Noisy Correspondence Makes CLIP More Robust
Haochen Han, Alex Jinpeng Wang, Peijun Ye +1
The data appetite for Vision-Language Models (VLMs) has continuously scaled up from the early millions to billions today, which faces an untenable trade-off with data quality and i…