Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling
Derek Li, Jiaming Zhou, Leo Maxime Brunswic +8
The pursuit of general-purpose artificial intelligence depends on large language models (LLMs) that can handle both structured reasoning and open-ended generation. We present Omni-…
cs.LG2024
HardCore Generation: Generating Hard UNSAT Problems for Data Augmentation
Joseph Cotnareanu, Zhanguang Zhang, Hui-Ling Zhen +2
Efficiently determining the satisfiability of a boolean equation -- known as the SAT problem for brevity -- is crucial in various industrial problems. Recently, the advent of deep…