2 papers
cs.CV2025
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
Jiawei Liu, Yuanzhi Zhu, Feiyu Gao +5
Generating visual text in natural scene images is a challenging task with many unsolved problems. Different from generating text on artificially designed images (such as posters, c…
cs.CL2024
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
Liang Chen, Zekun Wang, Shuhuai Ren +24
Building on the foundations of language modeling in natural language processing, Next Token Prediction (NTP) has evolved into a versatile training objective for machine learning ta…