2 papers
cs.LG2026
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning
Taojie Zhu, Dongyang Xu, Ding Zou +4
Post-training paradigms for Large Language Models (LLMs), primarily Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), face a fundamental dilemma: SFT provides stability…
cs.AI2026
Topology of Reasoning: Retrieved Cell Complex-Augmented Generation for Textual Graph Question Answering
Sen Zhao, Lincheng Zhou, Yue Chen +1
Retrieval-Augmented Generation (RAG) enhances the reasoning ability of Large Language Models (LLMs) by dynamically integrating external knowledge, thereby mitigating hallucinations…