6 papers · 1 filter
Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM
Xiaomeng Hu, Jiaqi Hu, Hao Chen +4
With the rapid advancement of large language models (LLMs), modern systems not only possess strong foundational capabilities and extensive knowledge, but can also solve complex pro…
SkillComposer: Learning to Evolve Agent Skills for Specification and Generalization
Qi Zhang, Zhaopeng Feng, Xiaonan Shi +8
Agent skills, which consist of reusable strategies that guide agent reasoning and action, have shown strong potential for improving model capability at inference time. However, cur…
DeltaMem: Towards Agentic Memory Management via Reinforcement Learning
Qi Zhang, Shen Huang, Chu Liu +4
Recent advances in persona-centric memory have revealed the powerful capability of multi-agent systems in managing persona memory, especially in conversational scenarios. However,…
LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization
Qi Zhang, Shouqing Yang, Lirong Gao +8
Large language models (LLMs) have demonstrated impressive capabilities in reasoning with the emergence of reasoning models like OpenAI-o1 and DeepSeek-R1. Recent research focuses o…
D.Va: Validate Your Demonstration First Before You Use It
Qi Zhang, Zhiqing Xiao, Ruixuan Xiao +2
In-context learning (ICL) has demonstrated significant potential in enhancing the capabilities of large language models (LLMs) during inference. It's well-established that ICL heav…
RECOST: External Knowledge Guided Data-efficient Instruction Tuning
Qi Zhang, Yiming Zhang, Haobo Wang +1
In the current landscape of large language models (LLMs), the process of instruction tuning serves as an essential step. Considering the high computing power overhead, data-efficie…