10 papers
Conformal Coverage Guarantees for Any Video Temporal Grounder
Aseel Mohamed, Rasul Khanbayov, Erchin Serpedin +1
Event boundaries in continuous video are ambiguous: re-annotate the same query-video pair and independent annotators mark moments that overlap by less than half on a large fraction…
Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading
Rasul Khanbayov, Hasan Kurban
Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic…
When Does Consensus Mean Correctness? Measuring the Agreement-Accuracy Coupling with Semantics-Preserving Re-Rendering
Rasul Khanbayov, Hasan Kurban
A model's agreement across perturbed inputs is used both as a label-free reliability signal and as a self-training target, on the premise that agreement tracks correctness. That co…
A Cost-Aware, Paired Protocol for Auditing Dynamic Tool Synthesis in Agentic Video Question Answering
Aseel Mohamed, Rama AlHamidi, Mohamed Rayan Barhdadi +3
Agentic Video Question Answering (VideoQA) systems invoke tools during inference, but their tool libraries are fixed, so recurring procedures are rebuilt from primitives on every q…
GRAPE: Graph-Augmented Prototype Explanations for Interactive Medical Image Diagnosis
Rasul Khanbayov, Erchin Serpedin, Hasan Kurban
Prototype-based medical image classifiers present three clinical limitations: they treat findings as independent, silently amplify unsafe physician feedback, and require full retra…
IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video
Rasul Khanbayov, Mohamed Rayan Barhdadi, Erchin Serpedin +1
Unsupervised physical parameter estimation from video lacks a common benchmark: existing methods evaluate on non-overlapping synthetic data, the sole real-world dataset is restrict…