5 papers · 1 filter
The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning
Aakriti Agrawal, Souradip Chakraborty, Armin Saghafian +6
Process Reward Models (PRMs) improve credit assignment for reasoning by providing step-level feedback. However, we identify a hidden bias in PRMs caused by severe imbalance in step…
Zero-shot Multivariate Time Series Forecasting Using Tabular Prior Fitted Networks
Mayuka Jayawardhana, Nihal Sharma, Kazem Meidani +3
Tabular foundation models, particularly Prior-data Fitted Networks like TabPFN have emerged as the leading contender in a myriad of tasks ranging from data imputation to label pred…
Learning Invariant Visual Representations for Planning with Joint-Embedding Predictive World Models
Leonardo F. Toso, Davit Shadunts, Yunyang Lu +4
World models learned from high-dimensional visual observations allow agents to make decisions and plan directly in latent space, avoiding pixel-level reconstruction. However, recen…
Bandits with Stochastic Experts: Constant Regret, Empirical Experts and Episodes
Nihal Sharma, Rajat Sen, Soumya Basu +2
We study a variant of the contextual bandit problem where an agent can intervene through a set of stochastic expert policies. Given a fixed context, each expert samples actions fro…
Bandits with Mean Bounds
Nihal Sharma, Soumya Basu, Karthikeyan Shanmugam +1
We study a variant of the bandit problem where side information in the form of bounds on the mean of each arm is provided. We prove that these translate to tighter estimates of sub…