collaborators

5 papers

cs.CL2026

JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation

Erlis Lushtaku, Bora Kargi, Ali Elganzory +4

LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their own code base, hardcode a spe…

cs.LG2026

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation

Bora Kargi, David Salinas

Evaluating new large language models typically requires costly human annotation campaigns at scale. LLM-as-a-judge offers a cheaper alternative, but judge scores carry systematic e…

cs.CV2026

Half-Truths Break Similarity-Based Retrieval

Bora Kargi, Arnas Uselis, Seong Joon Oh

When a text description is extended with an additional detail, image-text similarity should drop if that detail is wrong. We show that CLIP-style dual encoders often violate this i…

cs.CV2025

Evaluating Self-Supervised Learning in Medical Imaging: A Benchmark for Robustness, Generalizability, and Multi-Domain Impact

Valay Bundele, Karahan Sarıtaş, Bora Kargi +4

Self-supervised learning (SSL) has emerged as a promising paradigm in medical imaging, addressing the chronic challenge of limited labeled data in healthcare settings. While SSL ha…

cs.CL2025

Scholar Inbox: Personalized Paper Recommendations for Scientists

Markus Flicke, Glenn Angrabeit, Madhav Iyengar +10

Scholar Inbox is a new open-access platform designed to address the challenges researchers face in staying current with the rapidly expanding volume of scientific literature. We pr…