papers

Publications (8)

cs.CL2026

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +432

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…

cs.CL2026

Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models

Junru Lu, Jiarui Qin, Lingfeng Qiao +35

We introduce Youtu-LLM, a lightweight yet powerful language model that harmonizes high computational efficiency with native agentic intelligence. Unlike typical small models that r…

cs.CV2023

Coarse-to-Fine: Learning Compact Discriminative Representation for Single-Stage Image Retrieval

Yunquan Zhu, Xinkai Gao, Bo Ke +2

Image retrieval targets to find images from a database that are visually similar to the query image. Two-stage methods following retrieve-and-rerank paradigm have achieved excellen…

cs.CV2026

Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision

Zhixiang Wei, Yi Li, Zhehan Kan +38

Despite the significant advancements represented by Vision-Language Models (VLMs), current architectures often exhibit limitations in retaining fine-grained visual information, lea…

cs.CV2021

Head and Body: Unified Detector and Graph Network for Person Search in Media

Xiujun Shu, Yusheng Tao, Ruizhi Qiao +3

Person search in media has seen increasing potential in Internet applications, such as video clipping and character collection. This task is common but overlooked by previous perso…

cs.AI2026

From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents

Yifan Li, Shengbin Yue, Boyu Feng +6

The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks prioritize execution success, neglecting se…