2 papers
cs.CL2026
Learning When to Reason for Text-to-SQL via SFT and DPO
Soohyuk Jang, Jiheum Yeom, Nohil Park +4
Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmarks at the cost of high inferen…
cs.AI2026
Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering
Ik-hwan Kim, Hyeongrok Han, Mingi Jung +5
Large Language Models (LLMs) often produce incorrect answers on multi-hop question answering even when the reasoning trace already contains a correct intermediate conclusion. We at…