9 papers
Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets
Divake Kumar, Sina Tayebati, Devashri Naik +5
Computer-use agents turn vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejection, calibration, miss-severi…
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
Divake Kumar, Sina Tayebati, Devashri Naik +2
Vision-language models (VLMs) are increasingly used as automated judges for multimodal systems, yet their scores provide no indication of reliability. We study this problem through…
Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning
Ranganath Krishnan, Piyush Khanna, Omesh Tickoo
Large language models (LLMs) have revolutionized the field of natural language processing with their impressive reasoning and question-answering capabilities. However, these models…
TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic Reasoning
Sina Tayebati, Divake Kumar, Nastaran Darabi +3
Estimating uncertainty for AI agents in real-world multi-turn tool-using interaction with humans is difficult because failures are often triggered by sparse critical episodes (e.g.…
EigenTrack: Spectral Activation Feature Tracking for Hallucination and Out-of-Distribution Detection in LLMs and VLMs
Davide Ettori, Nastaran Darabi, Sina Tayebati +4
Large language models (LLMs) offer broad utility but remain prone to hallucination and out-of-distribution (OOD) errors. We propose EigenTrack, an interpretable real-time detector…
Learnable Conformal Prediction with Context-Aware Nonconformity Functions for Robotic Planning and Perception
Divake Kumar, Sina Tayebati, Francesco Migliarba +2
Deep learning models in robotics often output point estimates with poorly calibrated confidences, offering no native mechanism to quantify predictive reliability under novel, noisy…