7 papers · 1 filter
MInTRL: Off-policy Intervention can boost On-policy RL
Mingyu Chen, Yefan Tao, Gerald Friedland +2
Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the current policy but limiting learning to trajectories that the po…
Mitra-v2 Technical Report
Yefan Tao, Xiyuan Zhang, Xinyi Liu +13
We introduce Mitra-v2, a tabular foundation model that delivers state-of-the-art performance on real-world classification and regression problems, from credit-risk scoring and clin…
Cliff: Learning Process Rewards from the First Mistake
Peixuan Han, Runhui Wang, Ketan Ramaneti +3
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards le…
Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds
Yefan Tao, Gerald Friedland, Madhusudhanan Chandrasekaran +1
Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers. Large language models (LLMs) are increasingly prompted to "reflect,"…
Self-Aligned Reward: Towards Effective and Efficient Reasoners
Peixuan Han, Adit Krishnan, Gerald Friedland +2
Reinforcement learning with verifiable rewards has significantly advanced reasoning in large language models (LLMs), but such signals remain coarse, offering only binary correctnes…
Effects of Feature Correlations on Associative Memory Capacity
Stefan Bielmeier, Gerald Friedland
We investigate how feature correlations influence the capacity of Dense Associative Memory (DAM), a Transformer attention-like model. Practical machine learning scenarios involve f…