collaborators

55 papers

cs.LG2026

The Cost of Adaptivity: Matching Lower Bounds Across Learning Problems

Ibne Farabi Shihab, Adria Binte Habib

Adaptive procedures must work without nuisance information an oracle may use, such as a gradient scale or smoothness index, and robust procedures may have to answer queries whose c…

cs.LG2026

Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents

Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan

Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model distribution conditioned on a ha…

cs.LG2026

Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing

Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb

Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the…

cs.LG2026

Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify

Ibne Farabi Shihab, Fariya Afrin

Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increas…

cs.LG2026

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

Fariya Afrin, Ibne Farabi Shihab

Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this…

cs.AI2026

When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark…