collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning

Aakriti Agrawal, Souradip Chakraborty, Armin Saghafian +6

Process Reward Models (PRMs) improve credit assignment for reasoning by providing step-level feedback. However, we identify a hidden bias in PRMs caused by severe imbalance in step…

cs.LG2026

Zero-shot Multivariate Time Series Forecasting Using Tabular Prior Fitted Networks

Mayuka Jayawardhana, Nihal Sharma, Kazem Meidani +3

Tabular foundation models, particularly Prior-data Fitted Networks like TabPFN have emerged as the leading contender in a myriad of tasks ranging from data imputation to label pred…

cs.LG2026

Learning Invariant Visual Representations for Planning with Joint-Embedding Predictive World Models

Leonardo F. Toso, Davit Shadunts, Yunyang Lu +4

World models learned from high-dimensional visual observations allow agents to make decisions and plan directly in latent space, avoiding pixel-level reconstruction. However, recen…

cs.LG2024

Bandits with Stochastic Experts: Constant Regret, Empirical Experts and Episodes

Nihal Sharma, Rajat Sen, Soumya Basu +2

We study a variant of the contextual bandit problem where an agent can intervene through a set of stochastic expert policies. Given a fixed context, each expert samples actions fro…

cs.LG2024

Bandits with Mean Bounds

Nihal Sharma, Soumya Basu, Karthikeyan Shanmugam +1

We study a variant of the bandit problem where side information in the form of bounds on the mean of each arm is provided. We prove that these translate to tighter estimates of sub…