2 papers
cs.LG2026
WIST: Web-Grounded Iterative Self-Play Tree for Domain-Targeted Reasoning Improvement
Fangyuan Li, Pengfei Li, Shijie Wang +4
Recent progress in reinforcement learning with verifiable rewards (RLVR) offers a practical path to self-improvement of language models, but existing methods face a key trade-off:…
cs.LG2024
Fast and Slow Gradient Approximation for Binary Neural Network Optimization
Xinquan Chen, Junqi Gao, Biqing Qi +4
Binary Neural Networks (BNNs) have garnered significant attention due to their immense potential for deployment on edge devices. However, the non-differentiability of the quantizat…