3 papers
cs.CV2026
See the Text: From Tokenization to Visual Reading
Ling Xing, Rui Yan, Alex Jinpeng Wang +2
People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle ty…
cs.CV2026
Combating Noisy Labels through Fostering Self- and Neighbor-Consistency
Zeren Sun, Yazhou Yao, Tongliang Liu +3
Label noise is pervasive in various real-world scenarios, posing challenges in supervised deep learning. Deep networks are vulnerable to such label-corrupted samples due to the mem…
cs.CV2025
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
Henghao Zhao, Ge-Peng Ji, Rui Yan +2
The core challenge in video understanding lies in perceiving dynamic content changes over time. However, multimodal large language models struggle with temporal-sensitive video tas…