benchmark 1computer-use agents 1cross-platform evaluation 1reward modeling 1vision-language models 1
From the 1 of 12 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Reasoning Does Not Necessarily Improve Role-Playing Ability
Xiachong Feng, Longxu Dou, Lingpeng Kong
The application of role-playing large language models (LLMs) is rapidly expanding in both academic and commercial domains, driving an increasing demand for high-precision role-play…
cs.CL2025
A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios
Xiachong Feng, Longxu Dou, Ella Li +5
Game-theoretic scenarios have become pivotal in evaluating the social intelligence of Large Language Model (LLM)-based social agents. While numerous studies have explored these age…
cs.CL2024
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Chang Ma, Junlei Zhang, Zhihao Zhu +6
Evaluating Large Language Models (LLMs) as general-purpose agents is essential for understanding their capabilities and facilitating their integration into practical applications.…