4 papers · 1 filter
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
Yifeng Ding, Hung Le, Songyang Han +5
Training Large Language Models (LLMs) for multi-turn Tool-Integrated Reasoning (TIR) - where models iteratively reason, generate code, and verify through execution - remains challe…
Active Advantage-Aligned Online Reinforcement Learning with Offline Data
Xuefeng Liu, Hung T. C. Le, Siyu Chen +4
Online reinforcement learning (RL) enhances policies through direct interactions with the environment, but faces challenges related to sample efficiency. In contrast, offline RL le…
Reinforcement Learning for Causal Discovery without Acyclicity Constraints
Bao Duong, Hung Le, Biwei Huang +1
Recently, reinforcement learning (RL) has proved a promising alternative for conventional local heuristics in score-based approaches to learning directed acyclic causal graphs (DAG…
Iterative Multi-Agent Reinforcement Learning: A Novel Approach Toward Real-World Multi-Echelon Inventory Optimization
Georg Ziegner, Michael Choi, Hung Mac Chan Le +2
Multi-echelon inventory optimization (MEIO) is critical for effective supply chain management, but its inherent complexity can pose significant challenges. Heuristics are commonly…