activity
20212026
most citedContrastive Learning of User Behavior Sequence for Context-Aware Document Ranking

30 citations · 92 across the 12 of their papers we have counts for

collaborators

13 papers

cs.AI2026

OPOD: On-Policy Omni Distillation

Tong Zhao, Yuyang Hu, Yutao Zhu +5

Omni-modal models provide a unified interface for text, images, and audio. However, improving these abilities together remains difficult, as post-training on pooled multimodal data…

cs.AI2026

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

Rui Zou, Yutao Zhu, Mengqi Wei +1

Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challenging. Existing representativ…

cs.CL2024★ 2 cited

Towards Effective and Efficient Continual Pre-training of Large Language Models

Jie Chen, Zhipeng Chen, Jiapeng Wang +16

Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. To make the CPT approach more traceable, this paper presents…

cs.CL2024

YuLan: An Open-source Large Language Model

Yutao Zhu, Kun Zhou, Kelong Mao +35

Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many o…

cs.LG2024

An Integrated Data Processing Framework for Pretraining Foundation Models

Yiding Sun, Feng Wang, Yutao Zhu +2

The ability of the foundation models heavily relies on large-scale, diverse, and high-quality pretraining data. In order to improve data quality, researchers and practitioners ofte…

cs.IR2023

An Empirical Study of Uniform-Architecture Knowledge Distillation in Document Ranking

Xubo Qin, Xiyuan Liu, Xiongfeng Zheng +2

Although BERT-based ranking models have been commonly used in commercial search engines, they are usually time-consuming for online ranking tasks. Knowledge distillation, which aim…