8 papers · 1 filter
Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference
Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter +1
Static pruning imposes one sparse structure on every prompt, even though reasoning, retrieval, generation, coding, and translation can depend on different parts of a language model…
CS-WCP: Robust Conformal Sets for LLM-Judge Traffic Shifts with Uncertain Group Proportions
Ibne Farabi Shihab, Fariya Afrin
Prediction sets built from an LLM judge can undercover when deployment traffic changes the prevalence of task or policy groups. Weighted conformal prediction is exact under covaria…
When Accuracy Gaps Fail to Certify: Auditing Cross-Domain Recalibration of LLM Judges
Fariya Afrin, Ibne Farabi Shihab
A scalar recalibration map fitted for an LLM judge on one task can fail when the task distribution changes, but the source-target accuracy gap is often treated as a proxy for that…
Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify
Ibne Farabi Shihab, Fariya Afrin
Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increas…
Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees
Fariya Afrin, Ibne Farabi Shihab
Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this…
EST-PRM: Stress-Testing Process Reward Models Before They Become Load-Bearing
Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter +1
Process reward models (PRMs) are widely used in language-model training with dense step-level supervision. They assume PRM scores are stable proxies for step correctness under labe…