8 citations · 9 across the 7 of their papers we have counts for
6 papers · 1 filter
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context
Zhaowei Wang, Lishu Luo, Haodong Duan +9
Long-context modeling is becoming a core capability of modern large vision-language models (LVLMs), enabling sustained context management across long-document understanding, video…
Clair Obscur: an Illumination-Aware Method for Real-World Image Vectorization
Xingyue Lin, Shuai Peng, Xiangyu Xie +3
Image vectorization aims to convert raster images into editable, scalable vector representations while preserving visual fidelity. Existing vectorization methods struggle to repres…
Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression Recognition
Yu Li, Jin Jiang, Jianhua Zhu +4
Handwritten Mathematical Expression Recognition (HMER) remains a persistent challenge in Optical Character Recognition (OCR) due to the inherent freedom of symbol layouts and varia…
Vote&Mix: Plug-and-Play Token Reduction for Efficient Vision Transformer
Shuai Peng, Di Fu, Baole Wei +3
Despite the remarkable success of Vision Transformers (ViTs) in various visual tasks, they are often hindered by substantial computational cost. In this work, we introduce Vote\&Mi…
SketchRef: a Multi-Task Evaluation Benchmark for Sketch Synthesis
Xingyue Lin, Xingjian Hu, Shuai Peng +2
Sketching is a powerful artistic technique for capturing essential visual information about real-world objects and has increasingly attracted attention in image synthesis research.…
Handwritten Mathematical Expression Recognition with Bidirectionally Trained Transformer
Wenqi Zhao, Liangcai Gao, Zuoyu Yan +3
Encoder-decoder models have made great progress on handwritten mathematical expression recognition recently. However, it is still a challenge for existing methods to assign attenti…