11 papers
Imagining Recovery: Inference-Time Counterfactual Realignment for Vision-Language-Action Models
Yanyan Zhang, Disheng Liu, Kai Ye +6
Vision-language-action (VLA) models have improved the flexibility and generality of robotic manipulation, yet they remain fragile to online disruptions, such as changes in task goa…
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
Mohsen Hariri, Weicong Chen, Nahal Shahini +11
Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algori…
Medical Image Spatial Grounding with Semantic Sampling
Andrew Seohwan Yu, Mohsen Hariri, Kunio Nakamura +3
Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos. In medical imaging research, VLMs represent a bridge between object d…
CausalGuard: Conformal Inference under Graph Uncertainty
Vikash Singh, Weicong Chen, Debargha Ganguly +12
Estimating treatment effects from observational data requires choosing an adjustment set, but valid adjustment depends on an unknown causal graph. Graph misspecification can cause…
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
Mohsen Hariri, Amirhossein Samandar, Michael Hinczewski +1
Pass is widely used to report the reasoning performance of LLMs, but it often produces unstable and potentially misleading rankings, especially when the number of trials (sampl…
Scorio.jl: A Julia package for ranking stochastic responses
Mohsen Hariri, Michael Hinczewski, Vipin Chaudhary
Scorio.jl is a Julia package for evaluating and ranking systems from repeated responses to shared tasks. It provides a common tensor-based interface for direct score-based, pairwis…