activity
20242026
collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

MInTRL: Off-policy Intervention can boost On-policy RL

Mingyu Chen, Yefan Tao, Gerald Friedland +2

Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the current policy but limiting learning to trajectories that the po…

cs.LG2026

Mitra-v2 Technical Report

Yefan Tao, Xiyuan Zhang, Xinyi Liu +13

We introduce Mitra-v2, a tabular foundation model that delivers state-of-the-art performance on real-world classification and regression problems, from credit-risk scoring and clin…

cs.LG2026

Cliff: Learning Process Rewards from the First Mistake

Peixuan Han, Runhui Wang, Ketan Ramaneti +3

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards le…

cs.LG2026

Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds

Yefan Tao, Gerald Friedland, Madhusudhanan Chandrasekaran +1

Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers. Large language models (LLMs) are increasingly prompted to "reflect,"…

cs.LG2025

Self-Aligned Reward: Towards Effective and Efficient Reasoners

Peixuan Han, Adit Krishnan, Gerald Friedland +2

Reinforcement learning with verifiable rewards has significantly advanced reasoning in large language models (LLMs), but such signals remain coarse, offering only binary correctnes…

cs.LG2025

Effects of Feature Correlations on Associative Memory Capacity

Stefan Bielmeier, Gerald Friedland

We investigate how feature correlations influence the capacity of Dense Associative Memory (DAM), a Transformer attention-like model. Practical machine learning scenarios involve f…