43 papers
When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark…
Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation: Guarantees and Empirical Limits
Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma
Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate for federated, differentially private, adap…
CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning
Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan +2
Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can r…
Discrepancy-Rounded Fair Bandits with Static and Time-Varying Exposure Floors
Ibne Farabi Shihab, Joyanta Jyoti Mondal, Anuj Sharma
Minimum-exposure constraints arise in recommendation, content curation, and regulated allocation when each provider, arm, or group must receive guaranteed exposure inside a period…
Graph Dimensionality Reduction for Contextual Bandits: Structure-Specific Regret Bounds under Approximate Smoothness and Noisy Eigenspaces
Joyanta Jyoti Mondal, Ibne Farabi Shihab, Anuj Sharma
Contextual bandits with graph-structured arms arise in recommendation, citation retrieval, and social advertising, where arms connected on a graph tend to share reward signal. Stan…
EST-PRM: Stress-Testing Process Reward Models Before They Become Load-Bearing
Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter +1
Process reward models (PRMs) are widely used in language-model training with dense step-level supervision. They assume PRM scores are stable proxies for step correctness under labe…