5 papers
Allocate Marginal Reviews to Borderline Papers Using LLM Comparative Ranking
Elliot L. Epstein, Rajat Dwaraknath, John Winnicki +1
This paper argues that large ML conferences should allocate marginal review capacity primarily to papers near the acceptance boundary, rather than spreading extra reviews via rando…
LLMs are Overconfident: Evaluating Confidence Interval Calibration with FermiEval
Elliot L. Epstein, John Winnicki, Thanawat Sornwanee +1
Large language models (LLMs) excel at numerical estimation but struggle to correctly quantify uncertainty. We study how well LLMs construct confidence intervals around their own an…
Attention Factors for Statistical Arbitrage
Elliot L. Epstein, Rose Wang, Jaewon Choi +1
Statistical arbitrage exploits temporal price differences between similar assets. We develop a framework to jointly identify similar assets through factors, identify mispricing and…
SD-KDE: Score-Debiased Kernel Density Estimation
Elliot L. Epstein, Rajat Dwaraknath, Thanawat Sornwanee +2
We propose a novel method for density estimation that leverages an estimated score function to debias kernel density estimation (SD-KDE). In our approach, each data point is adjust…
MMMT-IF: A Challenging Multimodal Multi-Turn Instruction Following Benchmark
Elliot L. Epstein, Kaisheng Yao, Jing Li +2
Evaluating instruction following capabilities for multimodal, multi-turn dialogue is challenging. With potentially multiple instructions in the input model context, the task is tim…