3 papers
cs.CL2026
HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning
Yucan Guo, Xiaohan Wang, Miao Su +8
Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has be…
cs.CL2025
RouteRAG: Efficient Retrieval-Augmented Generation from Text and Graph via Reinforcement Learning
Yucan Guo, Miao Su, Saiping Guan +4
Retrieval-Augmented Generation (RAG) integrates non-parametric knowledge into Large Language Models (LLMs), typically from unstructured texts and structured graphs. While recent pr…
cs.AI2025
Mixture Policy based Multi-Hop Reasoning over N-tuple Temporal Knowledge Graphs
Zhongni Hou, Miao Su, Xiaolong Jin +4
Temporal Knowledge Graphs (TKGs), which utilize quadruples in the form of (subject, predicate, object, timestamp) to describe temporal facts, have attracted extensive attention. N-…