5 papers
TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization
Shichao Ma, Zhiyuan Ma, Ming Yang +8
Multi-turn tool-integrated reasoning enables Large Language Models (LLMs) to solve complex tasks through iterative information retrieval. However, current reinforcement learning (R…
TokenMixer-Large: Scaling Up Large Ranking Models in Industrial Recommenders
Yuchen Jiang, Jie Zhu, Xintian Han +18
While scaling laws for recommendation models have gained significant traction, existing architectures such as Wukong, HiFormer and DHEN, often struggle with sub-optimal designs and…
An item is worth one token in Multimodal Large Language Models-based Sequential Recommendation
Qiyong Zhong, Jiajie Su, Ming Yang +3
Sequential recommendations (SR) predict users' future interactions based on their historical behavior. The rise of Large Language Models (LLMs) has brought powerful generative and…
DART: Difficulty-Adaptive Reasoning Truncation for Efficient Large Language Models
Ruofan Zhang, Bin Xia, Zhen Cheng +4
Adaptive reasoning is essential for aligning the computational effort of large language models (LLMs) with the intrinsic difficulty of problems. Current chain-of-thought methods bo…
AROMA: Autonomous Rank-one Matrix Adaptation
Hao Nan Sheng, Zhi-yong Wang, Mingrui Yang +1
As large language models continue to grow in size, parameter-efficient fine-tuning (PEFT) has become increasingly crucial. While low-rank adaptation (LoRA) offers a solution throug…