Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
Yifeng Ding, Hung Le, Songyang Han +5
Training Large Language Models (LLMs) for multi-turn Tool-Integrated Reasoning (TIR) - where models iteratively reason, generate code, and verify through execution - remains challe…
cs.LG2025
Iterative Multi-Agent Reinforcement Learning: A Novel Approach Toward Real-World Multi-Echelon Inventory Optimization
Georg Ziegner, Michael Choi, Hung Mac Chan Le +2
Multi-echelon inventory optimization (MEIO) is critical for effective supply chain management, but its inherent complexity can pose significant challenges. Heuristics are commonly…
cs.LG2025
Active Advantage-Aligned Online Reinforcement Learning with Offline Data
Xuefeng Liu, Hung T. C. Le, Siyu Chen +4
Online reinforcement learning (RL) enhances policies through direct interactions with the environment, but faces challenges related to sample efficiency. In contrast, offline RL le…