Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Calibrate-Then-Delegate: Safety Monitoring with Risk and Budget Guarantees via Model Cascades
Edoardo Pona, Milad Kazemi, Mehran Hosseini +4
Monitoring LLM safety at scale requires balancing cost and accuracy: a cheap latent-space probe can screen every input, but hard cases should be escalated to a more expensive exper…
cs.LG2025
Abstract Counterfactuals for Language Model Agents
Edoardo Pona, Milad Kazemi, Yali Du +2
Counterfactual inference is a powerful tool for analysing and evaluating autonomous agents, but its application to language model (LM) agents remains challenging. Existing work on…