most citedTowards Agentic Recommender Systems in the Era of Multimodal Large Language Models

2 citations · 2 across the 4 of their papers we have counts for

collaborators

6 papers

cs.LG2025

Instance-Dependent Continuous-Time Reinforcement Learning via Maximum Likelihood Estimation

Runze Zhao, Yue Yu, Ruhan Wang +2

Continuous-time reinforcement learning (CTRL) provides a natural framework for sequential decision-making in dynamic environments where interactions evolve continuously over time.…

cs.LG2025

How to Provably Improve Return Conditioned Supervised Learning?

Zhishuai Liu, Yu Yang, Ruhan Wang +2

In sequential decision-making problems, Return-Conditioned Supervised Learning (RCSL) has gained increasing recognition for its simplicity and stability in modern decision-making t…

cs.LG2025

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality

Ruhan Wang, Zhiyong Wang, Chengkai Huang +5

For question-answering (QA) tasks, in-context learning (ICL) enables language models to generate responses without modifying their parameters by leveraging examples provided in the…

cs.LG2025

Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation

Runze Zhao, Yue Yu, Adams Yiyue Zhu +2

Continuous-time reinforcement learning (CTRL) provides a principled framework for sequential decision-making in environments where interactions evolve continuously over time. Despi…

cs.AI20252 cited

Towards Agentic Recommender Systems in the Era of Multimodal Large Language Models

Chengkai Huang, Junda Wu, Yu Xia +9

Recent breakthroughs in Large Language Models (LLMs) have led to the emergence of agentic AI systems that extend beyond the capabilities of standalone models. By empowering LLMs to…

cs.LG2025

Provable Zero-Shot Generalization in Offline Reinforcement Learning

Zhiyong Wang, Chen Yang, John C. S. Lui +1

In this work, we study offline reinforcement learning (RL) with zero-shot generalization property (ZSG), where the agent has access to an offline dataset including experiences from…