1 paper
Tianyu Liang, Xiangxi Zheng, Yilin Wang +1
Vision-Text Compression (VTC) renders long texts into images and encodes them through the vision encoder (ViT), compressing thousands of text tokens into far fewer visual tokens. H…