From the 1 of 5 linked papers with an AI index.
5 papers
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Zongxia Li, Zhongzhi Li, Yucheng Shi +10
The paper presents Long-Horizon-Terminal-Bench, a benchmark of 46 extended tasks with fine-grained intermediate rewards to evaluate AI agents' long-horizon planning and debugging a…
SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents
Tianming Sha, Yue Zhao, Lichao Sun +1
Autonomous AI agents can execute complex tasks with limited human review, yet they often lack the grounded operational knowledge to make their outputs not just executable but corre…
NavTrust: Benchmarking Trustworthiness for Embodied Navigation
Huaide Jiang, Yash Chaudhary, Yuping Wang +8
There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural language instructions; and Object-Goal Navigatio…
Benchmarking Vision Language Model Unlearning via Fictitious Facial Identity Dataset
Yingzi Ma, Jiongxiao Wang, Fei Wang +10
Machine unlearning has emerged as an effective strategy for forgetting specific information in the training data. However, with the increasing integration of visual data, privacy c…
Political-LLM: Large Language Models in Political Science
Lincan Li, Jiaqi Li, Catherine Chen +44
In recent years, large language models (LLMs) have been widely adopted in political science tasks such as election prediction, sentiment analysis, policy impact assessment, and mis…