From the 1 of 6 linked papers with an AI index.
6 papers
-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
Jiawei Xu, Minghui Liu, Juzheng Zhang +2
The paper proposes β‑OPSD, a generalized on‑policy self‑distillation method that treats the KL regularization weight as a tunable parameter, enabling a controlled interpolation bet…
The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning
Aakriti Agrawal, Souradip Chakraborty, Armin Saghafian +6
Process Reward Models (PRMs) improve credit assignment for reasoning by providing step-level feedback. However, we identify a hidden bias in PRMs caused by severe imbalance in step…
VeriGate: Verifier-Gated Step-Level Supervision for GRPO
Aakriti Agrawal, Minghui Liu, Furong Huang
Group Relative Policy Optimization (GRPO) is an effective recipe for training reasoning models with verifier-based outcome rewards, but its supervision is sparse: when all sampled…
Bridging the Divide: End-to-End Sequence-Graph Learning
Yuen Chen, Yulun Wu, Samuel Sharpe +5
Many real-world prediction tasks, particularly those involving entities such as customers or patients, involve both {sequential} and {relational} data. Each entity maintains its ow…
TimeSqueeze: Dynamic Patching for Efficient Time Series Forecasting
Sravan Kumar Ankireddy, Nikita Seleznev, Nam H. Nguyen +4
Transformer-based time series foundation models face a fundamental trade-off in choice of tokenization: point-wise embeddings preserve temporal fidelity but scale poorly with seque…
PersonaLedger: Generating Realistic Financial Transactions with Persona Conditioned LLMs and Rule Grounded Feedback
Dehao Yuan, Tyler Farnan, Stefan Tesliuc +8
Strict privacy regulations limit access to real transaction data, slowing open research in financial AI. Synthetic data can bridge this gap, but existing generators do not jointly…