7 papers
Aligning AI-driven discovery with human intuition
Kevin Zhang, Judah Goldfeder, Hod Lipson
As data-driven modeling of physical dynamical systems becomes more prevalent, a new challenge is emerging: making these models more compatible and aligned with existing human knowl…
CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning
Shijie Zhang, Zheng Xiao, Shiyu Liu +7
Online reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning abilities of large language models, but most methods still…
MOSAIC: Modular Orchestration for Structured Agentic Intelligence and Composition
Yifan Bao, Xinyu Xi, Xinyu Liu +8
Automated data science is a structured model-selection problem. A solution must choose data transformations, feature representations, architecture, training procedure, evaluation p…
Meta-probabilistic Modeling
Kevin Zhang, Yixin Wang
Probabilistic graphical models (PGMs) are widely used to discover latent structure in data, but their success hinges on selecting an appropriate model design. In practice, model sp…
Answer First, Reason Later: Aligning Search Relevance via Mode-Balanced Reinforcement Learning
Shijie Zhang, Xiang Guo, Rujun Guo +4
Building a search relevance model that achieves both low latency and high performance is a long-standing challenge in the search industry. To satisfy the millisecond-level response…
ETR: Outcome-Guided Elastic Trust Regions for Policy Optimization
Shijie Zhang, Kevin Zhang, Zheyuan Gu +5
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an important paradigm for unlocking reasoning capabilities in large language models, exemplified by the success…