2 papers
cs.LG2026
PLaID++: A Preference Aligned Language Model for Targeted Inorganic Materials Design
Andy Xu, Rohan Desai, Larry Wang +2
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising approach to improve correctness in LLMs, however, in many scientific problems, the objective is not…
cs.CL2025
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
Abdul Waheed, Chancharik Mitra, Laurie Z. Wang +2
Chain-of-thought reasoning, while powerful, can produce unnecessarily verbose output for simpler problems. We present a framework for difficulty-aware reasoning that teaches models…