#benchmark

topicbenchmark

86 papers · 1 filter

cs.AI2026

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

Kawai Chung, Chunkit Chan, Yauwai Yim +12

The paper introduces MultivationBench, a benchmark that tests multimodal large language models on their ability to reason about evolving human motivations across sequential visual…

cs.CL2026

ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models

Ruxi Gu, Zhenliang Zhang, Wei Wang

The paper introduces ForgetBench, a benchmark for measuring how large language models retain or forget factual and relational knowledge when they are continuously edited over time.

cs.SE2026

RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models

Yanlin Wang, Suiquan Wang, Yanli Wang +4

The paper presents RepoReasoner, a benchmark that evaluates how well large language models can reason about code across multiple files in a repository, testing both fine-grained ex…

cs.AI2026

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Chenrui Shi, Yuwei Wu, Yang Liu +5

The paper introduces an Interactive Reward Agent that evaluates GUI task completion by proposing conditions and verifying them using system, application, and GUI tools, and demonst…

cs.CV2026

CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking

Xiangqun Zhang, Likai Wang, Zekun Qian +2

The paper introduces CD-RMOT-Bench, a benchmark for evaluating how well referring multi-object tracking models trained on one visual domain perform on different, unseen domains, an…

cs.CV2026

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier +2

The paper presents Symbal, a dual‑stage method that uses off‑the‑shelf foundation models to automatically detect systematic misalignments—recurring caption errors tied to specific…