4 papers
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
Kaichen Zhang, Yuzhong Hong, Junwei Bao +4
Post-training plays a crucial role in refining and aligning large language models to meet specific tasks and human preferences. While recent advancements in post-training technique…
RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
Kaichen Zhang, Shenghao Gao, Yuzhong Hong +6
Current large language model post-training optimizes a risk-neutral objective that maximizes expected reward, yet evaluation relies heavily on risk-seeking metrics like Pass@k (at…
The Impact of Generative Artificial Intelligence on Market Equilibrium: Evidence from a Natural Experiment
Kaichen Zhang, Zixuan Yuan, Hui Xiong
Generative artificial intelligence (AI) exhibits the capability to generate creative content akin to human output with greater efficiency and reduced costs. This groundbreaking cap…
Optimized Cost Per Click in Online Advertising: A Theoretical Analysis
Kaichen Zhang, Zixuan Yuan, Hui Xiong
In recent years, Optimized Cost Per Click (OCPC) and Optimized Cost Per Mille (OCPM) have emerged as the most widely adopted pricing models in the online advertising industry. Howe…