2 papers
cs.LG2025
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
Chujie Zheng, Kai Dang, Bowen Yu +10
This paper proposes a novel formulation for reinforcement learning (RL) with large language models, explaining why and under what conditions the true sequence-level reward can be o…
cs.LG2025
TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks
Ziji Shi, Le Jiang, Ang Wang +6
Tensor parallelism is an essential technique for distributed training of large neural networks. However, automatically determining an optimal tensor parallel strategy is challengin…