3 papers
cs.LG2026
Inference-Time Code Selection via Symbolic Equivalence Partitioning
David Cho, Yifan Wang, Fanping Sui +1
Sampling multiple candidate programs at inference time is an effective way to improve LLM code generation. However, its benefit depends on reliably selecting a correct solution fro…
cs.AI2026
SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology
Yifan Wang, Bolian Li, David Cho +3
Reinforcement learning is critical to improving large reasoning models, but its success relies heavily on verifiable rewards (RLVR), making it hard to use in open-ended domains whe…
cs.AI2025
More is Less: The Pitfalls of Multi-Model Synthetic Preference Data in DPO Safety Alignment
Yifan Wang, Runjin Chen, Bolian Li +7
Aligning large language models (LLMs) with human values is an increasingly critical step in post-training. Direct Preference Optimization (DPO) has emerged as a simple, yet effecti…