55 papers
The Cost of Adaptivity: Matching Lower Bounds Across Learning Problems
Ibne Farabi Shihab, Adria Binte Habib
Adaptive procedures must work without nuisance information an oracle may use, such as a gradient scale or smoothness index, and robust procedures may have to answer queries whose c…
Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents
Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan
Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model distribution conditioned on a ha…
Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing
Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb
Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the…
Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify
Ibne Farabi Shihab, Fariya Afrin
Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increas…
Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees
Fariya Afrin, Ibne Farabi Shihab
Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this…
When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark…