4 papers
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
V Team, Wenyi Hong, Xiaotao Gu +94
We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depen…
Generative Compositor for Few-Shot Visual Information Extraction
Zhibo Yang, Wei Hua, Sibo Song +4
Visual Information Extraction (VIE), aiming at extracting structured information from visually rich document images, plays a pivotal role in document processing. Considering variou…
Visual Text Generation in the Wild
Yuanzhi Zhu, Jiawei Liu, Feiyu Gao +6
Recently, with the rapid advancements of generative models, the field of visual text generation has witnessed significant progress. However, it is still challenging to render high-…
HIP: Hierarchical Point Modeling and Pre-training for Visual Information Extraction
Rujiao Long, Pengfei Wang, Zhibo Yang +1
End-to-end visual information extraction (VIE) aims at integrating the hierarchical subtasks of VIE, including text spotting, word grouping, and entity labeling, into a unified fra…