5 papers · 1 filter
Recursive Harness Self-Improvement
Hyunin Lee, Jinglue Xu, Jeffrey Seely +3
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This…
Sakana Fugu Technical Report
Yujin Tang, Edoardo Cetin, Jinglue Xu +11
The capabilities of frontier Large Language Models (LLMs) continue to advance, with different providers increasingly specializing in distinct domains. This raises a natural next ob…
Cross-attention Secretly Performs Orthogonal Alignment in Recommendation Models
Hyunin Lee, Yong Zhang, Hoang Vu Nguyen +8
Cross-domain sequential recommendation (CDSR) aims to align heterogeneous user behavior sequences collected from different domains. While cross-attention is widely used to enhance…
Beyond Exact Gradients: Convergence of Stochastic Soft-Max Policy Gradient Methods with Entropy Regularization
Yuhao Ding, Junzi Zhang, Hyunin Lee +1
Entropy regularization is an efficient technique for encouraging exploration and preventing a premature convergence of (vanilla) policy gradient methods in reinforcement learning (…
Pausing Policy Learning in Non-stationary Reinforcement Learning
Hyunin Lee, Ming Jin, Javad Lavaei +1
Real-time inference is a challenge of real-world reinforcement learning due to temporal differences in time-varying environments: the system collects data from the past, updates th…