activity
20242026
most citedDetecting Proxy Gaming in RL and LLM Alignment via Evaluator Stress Tests

1 citations · 1 across the 27 of their papers we have counts for

collaborators
Showing cs.LGShow all

28 papers · 1 filter

cs.LG2026

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan +2

Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can r…

cs.LG2026

Discrepancy-Rounded Fair Bandits with Static and Time-Varying Exposure Floors

Ibne Farabi Shihab, Joyanta Jyoti Mondal, Anuj Sharma

Minimum-exposure constraints arise in recommendation, content curation, and regulated allocation when each provider, arm, or group must receive guaranteed exposure inside a period…

cs.LG2026

Graph Dimensionality Reduction for Contextual Bandits: Structure-Specific Regret Bounds under Approximate Smoothness and Noisy Eigenspaces

Joyanta Jyoti Mondal, Ibne Farabi Shihab, Anuj Sharma

Contextual bandits with graph-structured arms arise in recommendation, citation retrieval, and social advertising, where arms connected on a graph tend to share reward signal. Stan…

cs.LG2026

EST-PRM: Stress-Testing Process Reward Models Before They Become Load-Bearing

Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter +1

Process reward models (PRMs) are widely used in language-model training with dense step-level supervision. They assume PRM scores are stable proxies for step correctness under labe…

cs.LG2026

Grounded Decoding: Retrieval-Anchored Probability Fusion for Faithful RAG

Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter +1

As retrieval-augmented generation (RAG) systems scale, it becomes increasingly challenging to ensure faithful grounding in external evidence. Large language models may still priori…

cs.LG2026

Topology-Aware State Abstraction with Tangle Cores for Markov Decision Processes

Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

State abstraction in reinforcement learning is usually formulated as a partition of states based on reward and transition similarity. This excludes a common structural pattern in n…