collaborators

5 papers

cs.LG2026

Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents

Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan

Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model distribution conditioned on a ha…

cs.LG2026

Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing

Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb

Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the…

cs.LG2026

Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb

Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare because probability mode,…

cs.LG2026

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan +2

Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can r…

cs.LG2026

Canonicalized Stable-List Replay for Private Federated Continual Learning over Language-Model Embeddings

Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Anuj Sharma

Federated continual learning (FCL) lets distributed clients adapt language-model heads to evolving NLP tasks without sharing raw text. Under user-level differential privacy (DP), r…