2 papers
cs.CL2026
SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning
Tao Liu, Tao Feng, Xiangheng Li +9
Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods use execution correctness onl…
cs.CL2026
KbSD: Knowledge Boundary aware Self-Distillation for Behavioral Calibration in Agentic Search
Tao Feng, Xinke Jiang, Chao Wu
Agentic search equips large language models with dynamic retrieval abilities, but existing reinforcement learning methods remain limited by reward sparsity in knowledge boundary ca…