2 papers
cs.AI2026
PROOF-Gen: From Optimized Data to Better Distillation
Anh Ta, Junjie Zhu, Shahin Shayandeh
Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that d…
cs.AI2026
Reinforced Agent: Inference-Time Feedback for Tool-Calling Agents
Anh Ta, Junjie Zhu, Shahin Shayandeh
Tool-calling agents are evaluated on tool selection, parameter accuracy, and scope recognition, yet LLM trajectory assessments remain inherently post-hoc. Disconnected from the act…