5 papers
Medical Image Spatial Grounding with Semantic Sampling
Andrew Seohwan Yu, Mohsen Hariri, Kunio Nakamura +3
Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos. In medical imaging research, VLMs represent a bridge between object d…
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
Mohsen Hariri, Amirhossein Samandar, Michael Hinczewski +1
Pass is widely used to report the reasoning performance of LLMs, but it often produces unstable and potentially misleading rankings, especially when the number of trials (sampl…
LRD-Net: A Lightweight Real-Centered Detection Network for Cross-Domain Face Forgery Detection
Xuecen Zhang, Vipin Chaudhary
The rapid advancement of diffusion-based generative models has made face forgery detection a critical challenge in digital forensics. Current detection methods face two fundamental…
Scorio.jl: A Julia package for ranking stochastic responses
Mohsen Hariri, Michael Hinczewski, Vipin Chaudhary
Scorio.jl is a Julia package for evaluating and ranking systems from repeated responses to shared tasks. It provides a common tensor-based interface for direct score-based, pairwis…
Rethinking Vision Transformer Depth via Structural Reparameterization
Chengwei Zhou, Vipin Chaudhary, Gourav Datta
The computational overhead of Vision Transformers in practice stems fundamentally from their deep architectures, yet existing acceleration strategies have primarily targeted algori…