activity
20222024
most citedBOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents

9 citations · 18 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20241 cited

Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Ruohong Zhang, Liangke Gui, Zhiqing Sun +8

Preference modeling techniques, such as direct preference optimization (DPO), has shown effective in enhancing the generalization abilities of large language model (LLM). However,…

cs.CL20241 cited

FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability

Congying Xia, Chen Xing, Jiangshu Du +5

This paper presents FoFo, a pioneering benchmark for evaluating large language models' (LLMs) ability to follow complex, domain-specific formats, a crucial yet underexamined capabi…

cs.AI20239 cited

BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents

Zhiwei Liu, Weiran Yao, Jianguo Zhang +12

The massive successes of large language models (LLMs) encourage the emerging exploration of LLM-augmented Autonomous Agents (LAAs). An LAA is able to generate actions with its core…

cs.CL20236 cited

Fantastic Rewards and How to Tame Them: A Case Study on Reward Learning for Task-oriented Dialogue Systems

Yihao Feng, Shentao Yang, Shujian Zhang +4

When learning task-oriented dialogue (ToD) agents, reinforcement learning (RL) techniques can naturally be utilized to train dialogue strategies to achieve user-specific goals. Pri…

cs.LG20221 cited

Operator Deep Q-Learning: Zero-Shot Reward Transferring in Reinforcement Learning

Ziyang Tang, Yihao Feng, Qiang Liu

Reinforcement learning (RL) has drawn increasing interests in recent years due to its tremendous success in various applications. However, standard RL algorithms can only be applie…