4 papers
CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning
Ke Sun, Yizhou Zhao, Jiayi Xin +2
Context or prompt-level reweighting has emerged as a central algorithmic lever in Reinforcement Learning with Verified Rewards (RLVR) for improving the reasoning capability of larg…
Memory Bear AI A Breakthrough from Memory to Cognition Toward Artificial General Intelligence
Deliang Wen, Ke Sun
Large language models (LLMs) face inherent limitations in memory, including restricted context windows, long-term knowledge forgetting, redundant information accumulation, and hall…
Knowledge Graph Tokenization for Behavior-Aware Generative Next POI Recommendation
Ke Sun, Mayi Xu
Generative paradigm, especially powered by Large Language Models (LLMs), has emerged as a new solution to the next point-of-interest (POI) recommendation. Pioneering studies usuall…
Mistral-C2F: Coarse to Fine Actor for Analytical and Reasoning Enhancement in RLHF and Effective-Merged LLMs
Chen Zheng, Ke Sun, Xun Zhou
Despite the advances in Large Language Models (LLMs), exemplified by models like GPT-4 and Claude, smaller-scale LLMs such as Llama and Mistral often struggle with generating in-de…