18 papers
PageLLM: A Multi-Grained Reward Framework for Whole-Page Optimization with Large Language Models
Xinyuan Wang, Liang Wu, Dongjie Wang +1
Whole-page optimization (WPO) decides how search and recommendation results are surfaced to users, and large language models (LLMs) open a new route to it by treating page generati…
Causally-Guided Diffusion for Stable Feature Selection
Arun Vignesh Malarkkan, Xinyuan Wang, Kunpeng Liu +2
Feature selection is fundamental to robust data-centric AI, but most existing methods optimize predictive performance under a single data distribution. This often selects spurious…
BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning
Yuan Li, Bo Wang, Yufei Gao +4
Proximal constraints are fundamental to the stability of the Large Language Model reinforcement learning. While the canonical clipping mechanism in PPO serves as an efficient surro…
Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation
Xinyuan Wang, Kunpeng Liu, Arun Vignesh Malarkkan +1
Feature Transformation (FT) is a core data-centric AI task that improves feature space quality to advance downstream predictive performance. However, discovering effective transfor…
Efficient Post-Training Refinement of Latent Reasoning in Large Language Models
Xinyuan Wang, Dongjie Wang, Wangyang Ying +5
Reasoning is a key component of language understanding in Large Language Models. While Chain-of-Thought prompting enhances performance via explicit intermediate steps, it suffers f…
Data-Efficient Symbolic Regression via Foundation Model Distillation
Wangyang Ying, Jinghan Zhang, Haoyue Bai +5
Discovering interpretable mathematical equations from observed data (a.k.a. equation discovery or symbolic regression) is a cornerstone of scientific discovery, enabling transparen…