activity
20232025
most citedRobust Multi-agent Communication via Multi-view Message Certification

40 citations · 41 across the 6 of their papers we have counts for

collaborators

7 papers

cs.MA2025

Multi-agent In-context Coordination via Decentralized Memory Retrieval

Tao Jiang, Zichuan Lin, Lihe Li +6

Large transformer models, trained on diverse datasets, have demonstrated impressive few-shot performance on previously unseen tasks without requiring parameter updates. This capabi…

cs.CL2025

Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference

Wenjie Qiu, Yi-Chen Li, Xuqin Zhang +4

Learning reward models from human preference datasets and subsequently optimizing language models via reinforcement learning has emerged as a fundamental paradigm for aligning LLMs…

cs.LG2024

Stable Continual Reinforcement Learning via Diffusion-based Trajectory Replay

Feng Chen, Fuguang Han, Cong Guan +4

Given the inherent non-stationarity prevalent in real-world applications, continual Reinforcement Learning (RL) aims to equip the agent with the capability to address a series of s…

cs.LG2024

Hindsight Preference Learning for Offline Preference-based Reinforcement Learning

Chen-Xiao Gao, Shengjun Fang, Chenjun Xiao +2

Offline preference-based reinforcement learning (RL), which focuses on optimizing policies using human preferences between pairs of trajectory segments selected from an offline dat…

cs.CL2024

Improving Sample Efficiency of Reinforcement Learning with Background Knowledge from Large Language Models

Fuxiang Zhang, Junyou Li, Yi-Chen Li +3

Low sample efficiency is an enduring challenge of reinforcement learning (RL). With the advent of versatile large language models (LLMs), recent works impart common-sense knowledge…

cs.LG2024★ 1 cited

Q-Adapter: Customizing Pre-trained LLMs to New Preferences with Forgetting Mitigation

Yi-Chen Li, Fuxiang Zhang, Wenjie Qiu +5

Large Language Models (LLMs), trained on a large amount of corpus, have demonstrated remarkable abilities. However, it may not be sufficient to directly apply open-source LLMs like…