Publications (8)
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
Junru Lu, Jiarui Qin, Lingfeng Qiao +35
We introduce Youtu-LLM, a lightweight yet powerful language model that harmonizes high computational efficiency with native agentic intelligence. Unlike typical small models that r…
Coarse-to-Fine: Learning Compact Discriminative Representation for Single-Stage Image Retrieval
Yunquan Zhu, Xinkai Gao, Bo Ke +2
Image retrieval targets to find images from a database that are visually similar to the query image. Two-stage methods following retrieve-and-rerank paradigm have achieved excellen…
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
Zhixiang Wei, Yi Li, Zhehan Kan +38
Despite the significant advancements represented by Vision-Language Models (VLMs), current architectures often exhibit limitations in retaining fine-grained visual information, lea…
Head and Body: Unified Detector and Graph Network for Person Search in Media
Xiujun Shu, Yusheng Tao, Ruizhi Qiao +3
Person search in media has seen increasing potential in Internet applications, such as video clipping and character collection. This task is common but overlooked by previous perso…
From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
Yifan Li, Shengbin Yue, Boyu Feng +6
The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks prioritize execution success, neglecting se…