5 papers
Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs
Shrutendra Harsola, Vignesh Subrahmaniam, Vikas Raturi +7
Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business changes rather than randomized…
RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning
Xiang Gao, Yuguang Yao, Qi Zhang +5
Large language models (LLMs) often struggle to use tools reliably in domain-specific settings, where APIs may be idiosyncratic, under-documented, or tailored to private workflows.…
When Evidence is Sparse: Weakly Supervised Early Failure Alerting in Dialogs and LLM-Agent Trajectories
Avinash Baidya, Xinran Liang, Ruocheng Guo +2
Early failure alerting requires deciding, while a dialog or agent trajectory is still unfolding, whether to flag it as likely to fail. This is challenging because supervision is ty…
Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
Ruocheng Guo, Kaiwen Dong, Xiang Gao +1
While most efforts to improve LLM-based tool-using agents focus on the agent itself - through larger models, better prompting, or fine-tuning - agent performance increasingly plate…
Node-Level Uncertainty Estimation in LLM-Generated SQL
Hilaf Hasson, Ruocheng Guo
We present a practical framework for detecting errors in LLM-generated SQL by estimating uncertainty at the level of individual nodes in the query's abstract syntax tree (AST). Our…