3 papers
cs.CV2026
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
Zhixiang Wei, Yi Li, Zhehan Kan +38
Despite the significant advancements represented by Vision-Language Models (VLMs), current architectures often exhibit limitations in retaining fine-grained visual information, lea…
cs.CV2026
Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding
Haoyu Cao, Kun Yin, Yunfei Wu +16
This paper presents Youtu-Parsing, an efficient and versatile document parsing model designed for high-performance content extraction. The architecture employs a native Vision Tran…
cs.CL2024
Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and Reaction
Haoqiu Yan, Yongxin Zhu, Kai Zheng +4
Large Language Model (LLM)-enhanced agents become increasingly prevalent in Human-AI communication, offering vast potential from entertainment to professional domains. However, cur…