activity
20242026
collaborators

8 papers

cs.CL2026

Hint-Guided Diversified Policy Optimization for LLM Reasoning

Zhiyu Cao, Kaixin Wu, Mingjie Zhong +4

Recent developments in Large Language Models (LLMs) have showcased impressive reasoning capabilities, with Reinforcement Learning with Verifiable Rewards (RLVR) being a promising e…

cs.LG2026

GAST: Gradient-aligned Sparse Tuning of Large Language Models with Data-layer Selection

Kai Yao, Zhenghan Song, Kaixin Wu +5

Parameter-Efficient Fine-Tuning (PEFT) has become a key strategy for adapting large language models, with recent advances in sparse tuning reducing overhead by selectively updating…

cs.CL2025

GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models

Kai Yao, Zhaorui Tan, Penglei Gao +7

The rapid growth of large language models (LLMs) with traditional centralized fine-tuning emerges as a key technique for adapting these models to domain-specific challenges, yieldi…

cs.AI2025

A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning

Yixin Ji, Juntao Li, Yang Xiang +6

The remarkable performance of the o1 model in complex reasoning demonstrates that test-time compute scaling can further unlock the model's potential, enabling powerful System-2 thi…

cs.IR2025

Alleviating LLM-based Generative Retrieval Hallucination in Alipay Search

Yedan Shen, Kaixin Wu, Yuechen Ding +6

Generative retrieval (GR) has revolutionized document retrieval with the advent of large language models (LLMs), and LLM-based GR is gradually being adopted by the industry. Despit…

cs.AI2025

CPRM: A LLM-based Continual Pre-training Framework for Relevance Modeling in Commercial Search

Kaixin Wu, Yixin Ji, Zeyuan Chen +9

Relevance modeling between queries and items stands as a pivotal component in commercial search engines, directly affecting the user experience. Given the remarkable achievements o…