1 paper · 1 filter
Shenglai Zeng, Qirui Wang, Kai Guo +3
Visual Text Comprehension (VTC) renders text into images for a vision-language model (VLM) to read, sidestepping LLM context-window limits and powering applications from long-page…