activity
20222025
most citedAMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

6 citations · 12 across the 13 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2025

One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning

Yuan Pu, Yazhe Niu, Jia Tang +3

In heterogeneous multi-task decision-making, tasks not only exhibit diverse observation and action spaces but also vary substantially in their underlying complexities. While conven…

cs.LG2025

Empowering LLMs in Decision Games through Algorithmic Data Synthesis

Haolin Wang, Xueyan Li, Yazhe Niu +2

Large Language Models (LLMs) have exhibited impressive capabilities across numerous domains, yet they often struggle with complex reasoning and decision-making tasks. Decision-maki…

cs.LG2024

Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective

Jinouwen Zhang, Rongkun Xue, Yazhe Niu +4

Generative models, particularly diffusion models, have achieved remarkable success in density estimation for multimodal data, drawing significant interest from the reinforcement le…

cs.LG2024

UniZero: Generalized and Efficient Planning with Scalable Latent World Models

Yuan Pu, Yazhe Niu, Zhenjie Yang +3

Learning predictive world models is crucial for enhancing the planning capabilities of reinforcement learning (RL) agents. Recently, MuZero-style algorithms, leveraging the value e…

cs.LG2023

A Perspective of Q-value Estimation on Offline-to-Online Reinforcement Learning

Yinmin Zhang, Jie Liu, Chuming Li +4

Offline-to-online Reinforcement Learning (O2O RL) aims to improve the performance of offline pretrained policy using only a few online samples. Built on offline RL algorithms, most…

cs.LG2023★ 2 cited

LightZero: A Unified Benchmark for Monte Carlo Tree Search in General Sequential Decision Scenarios

Yazhe Niu, Yuan Pu, Zhenjie Yang +6

Building agents based on tree-search planning capabilities with learned models has achieved remarkable success in classic decision-making problems, such as Go and Atari. However, i…