4 papers
Calibrate-Then-Delegate: Safety Monitoring with Risk and Budget Guarantees via Model Cascades
Edoardo Pona, Milad Kazemi, Mehran Hosseini +4
Monitoring LLM safety at scale requires balancing cost and accuracy: a cheap latent-space probe can screen every input, but hard cases should be escalated to a more expensive exper…
LTL Verification of Memoryful Neural Agents
Mehran Hosseini, Alessio Lomuscio, Nicola Paoletti
We present a framework for verifying Memoryful Neural Multi-Agent Systems (MN-MAS) against full Linear Temporal Logic (LTL) specifications. In MN-MAS, agents interact with a non-de…
Certified Guidance for Planning with Deep Generative Models
Francesco Giacomarra, Mehran Hosseini, Nicola Paoletti +1
Deep generative models, such as generative adversarial networks and diffusion models, have recently emerged as powerful tools for planning tasks and behavior synthesis in autonomou…
Verifiably Robust Conformal Prediction
Linus Jeary, Tom Kuipers, Mehran Hosseini +1
Conformal Prediction (CP) is a popular uncertainty quantification method that provides distribution-free, statistically valid prediction sets, assuming that training and test data…