works on

From the 1 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

cs.AI2026

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Zongxia Li, Zhongzhi Li, Yucheng Shi +10

The paper presents Long-Horizon-Terminal-Bench, a benchmark of 46 extended tasks with fine-grained intermediate rewards to evaluate AI agents' long-horizon planning and debugging a…

cs.AI2026

SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents

Tianming Sha, Yue Zhao, Lichao Sun +1

Autonomous AI agents can execute complex tasks with limited human review, yet they often lack the grounded operational knowledge to make their outputs not just executable but corre…

cs.RO2026

NavTrust: Benchmarking Trustworthiness for Embodied Navigation

Huaide Jiang, Yash Chaudhary, Yuping Wang +8

There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural language instructions; and Object-Goal Navigatio…

cs.CV2025

Benchmarking Vision Language Model Unlearning via Fictitious Facial Identity Dataset

Yingzi Ma, Jiongxiao Wang, Fei Wang +10

Machine unlearning has emerged as an effective strategy for forgetting specific information in the training data. However, with the increasing integration of visual data, privacy c…

cs.CL2024

Political-LLM: Large Language Models in Political Science

Lincan Li, Jiaqi Li, Catherine Chen +44

In recent years, large language models (LLMs) have been widely adopted in political science tasks such as election prediction, sentiment analysis, policy impact assessment, and mis…