18 papers
Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking
Chia-Hsuan Lee, Sihui Dai, Mingyang Zhou +5
Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction that consume tokens without im…
SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision
Chia-Hsuan Lee, Zelei Cheng, Yu Wang +4
On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence. Incoherent rollouts yield noisy gradie…
Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
Sangwoo Cho, Kushal Chawla, Pengshan Cai +4
Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, an…
SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation
Chenyang Zhu, Jiayu Yao, Kushal Chawla +10
As autonomous agents tackle increasingly complex multi-step, multi-agent tasks, their execution trajectories have scaled beyond the constraints of even the largest context windows.…
Towards Scalable Customization and Deployment of Multi-Agent Systems for Enterprise Applications
Paresh Dashore, Shreyas Kulkarni, Uttam Gurram +5
Large language model (LLM)-based multi-agent systems demonstrate strong performance on complex reasoning and task execution, enabling broad enterprise applications. However, produc…
T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains
Genta Indra Winata, Amartya Chakraborty, Yuzhen Lin +12
Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However, existing benchmarks remain li…