8 papers
Hint-Guided Diversified Policy Optimization for LLM Reasoning
Zhiyu Cao, Kaixin Wu, Mingjie Zhong +4
Recent developments in Large Language Models (LLMs) have showcased impressive reasoning capabilities, with Reinforcement Learning with Verifiable Rewards (RLVR) being a promising e…
GAST: Gradient-aligned Sparse Tuning of Large Language Models with Data-layer Selection
Kai Yao, Zhenghan Song, Kaixin Wu +5
Parameter-Efficient Fine-Tuning (PEFT) has become a key strategy for adapting large language models, with recent advances in sparse tuning reducing overhead by selectively updating…
GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models
Kai Yao, Zhaorui Tan, Penglei Gao +7
The rapid growth of large language models (LLMs) with traditional centralized fine-tuning emerges as a key technique for adapting these models to domain-specific challenges, yieldi…
A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning
Yixin Ji, Juntao Li, Yang Xiang +6
The remarkable performance of the o1 model in complex reasoning demonstrates that test-time compute scaling can further unlock the model's potential, enabling powerful System-2 thi…
Alleviating LLM-based Generative Retrieval Hallucination in Alipay Search
Yedan Shen, Kaixin Wu, Yuechen Ding +6
Generative retrieval (GR) has revolutionized document retrieval with the advent of large language models (LLMs), and LLM-based GR is gradually being adopted by the industry. Despit…
CPRM: A LLM-based Continual Pre-training Framework for Relevance Modeling in Commercial Search
Kaixin Wu, Yixin Ji, Zeyuan Chen +9
Relevance modeling between queries and items stands as a pivotal component in commercial search engines, directly affecting the user experience. Given the remarkable achievements o…