3 papers
cs.LG2026
Prioritized Replay for RL Post-training
Mehdi Fatemi
We introduce a problem-level prioritization framework for RL post-training of large language models. Building on insights from prioritized replay in deep RL, as well as prior obser…
cs.CL2025
Concise Reasoning via Reinforcement Learning
Mehdi Fatemi, Banafsheh Rafiee, Mingjie Tang +1
A major drawback of reasoning models is their excessive token usage, inflating computational cost, resource demand, and latency. We show this verbosity stems not from deeper reason…
cs.CL2024
OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset
Allen Roush, Yusuf Shabazz, Arvind Balaji +7
We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community. This dataset includes over 3.…