Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
Zhiyang Xu, Minqian Liu, Ying Shen +5
Recent advancements in Vision-Language Models (VLMs) have led to the emergence of Vision-Language Generalists (VLGs) capable of understanding and generating both text and images. H…
cs.CL2024
RE: Region-Aware Relation Extraction from Visually Rich Documents
Pritika Ramu, Sijia Wang, Lalla Mouatadid +2
Current research in form understanding predominantly relies on large pre-trained language models, necessitating extensive data for pre-training. However, the importance of layout s…