activity
20232026
collaborators

7 papers

cs.LG2026

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation

Sijia Li, Yuchen Huang, Zifan Liu +7

Reinforcement learning has become a widely used post-training approach for LLM agents, where training commonly relies on outcome-level rewards that provide only coarse supervision.…

cs.LG2026

GAS: Enhancing Reward-Cost Balance of Generative Model-assisted Offline Safe RL

Zifan Liu, Xinran Li, Shibo Chen +1

Offline Safe Reinforcement Learning (OSRL) aims to learn a policy to achieve high performance in sequential decision-making while satisfying constraints, using only pre-collected d…

cs.LG2025

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory

Sijia Li, Yuchen Huang, Zifan Liu +6

As intents unfold and environments change, multi-turn agents face continuously shifting decision contexts. Although reusing past experience is intuitively appealing, existing appro…

cs.LG2024

TSDS: Data Selection for Task-Specific Model Finetuning

Zifan Liu, Amin Karbasi, Theodoros Rekatsinas

Finetuning foundation models for specific tasks is an emerging paradigm in modern machine learning. The efficacy of task-specific finetuning largely depends on the selection of app…

cs.LG2024

Reinforcement Learning with Intrinsically Motivated Feedback Graph for Lost-sales Inventory Control

Zifan Liu, Xinran Li, Shibo Chen +3

Reinforcement learning (RL) has proven to be well-performed and general-purpose in the inventory control (IC). However, further improvement of RL algorithms in the IC domain is imp…

cs.LG2024

Individual Contributions as Intrinsic Exploration Scaffolds for Multi-agent Reinforcement Learning

Xinran Li, Zifan Liu, Shibo Chen +1

In multi-agent reinforcement learning (MARL), effective exploration is critical, especially in sparse reward environments. Although introducing global intrinsic rewards can foster…