5 papers
Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify
Ibne Farabi Shihab, Fariya Afrin
Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increas…
Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees
Fariya Afrin, Ibne Farabi Shihab
Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this…
EST-PRM: Stress-Testing Process Reward Models Before They Become Load-Bearing
Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter +1
Process reward models (PRMs) are widely used in language-model training with dense step-level supervision. They assume PRM scores are stable proxies for step correctness under labe…
Grounded Decoding: Retrieval-Anchored Probability Fusion for Faithful RAG
Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter +1
As retrieval-augmented generation (RAG) systems scale, it becomes increasingly challenging to ensure faithful grounding in external evidence. Large language models may still priori…
Dynamic Proxy-Mixing: Transferring Replay Controllers from Small to Large Models for Continual Instruction Tuning
Ibne Farabi Shihab, Fariya Afrin, Anuj Sharma
Continual instruction tuning updates a language model through a sequence of new domains, yet each update can progressively erode previously learned capabilities and alignment behav…