2 papers
cs.LG2026
Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models
Matthew DosSantos DiSorbo, Harang Ju
Effective automation hinges on deciding when to act and when to escalate. We model this as a decision under uncertainty: an LLM forms a prediction, estimates its probability of bei…
cs.AI2026
Teaching AI to Handle Exceptions: Supervised Fine-Tuning with Human-Aligned Judgment
Matthew DosSantos DiSorbo, Harang Ju, Sinan Aral
Large language models (LLMs), initially developed for generative AI, are now evolving into agentic AI systems, which make decisions in complex, real-world contexts. Unfortunately,…