#benchmark
86 papers · 1 filter
MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
Kawai Chung, Chunkit Chan, Yauwai Yim +12
The paper introduces MultivationBench, a benchmark that tests multimodal large language models on their ability to reason about evolving human motivations across sequential visual…
ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models
Ruxi Gu, Zhenliang Zhang, Wei Wang
The paper introduces ForgetBench, a benchmark for measuring how large language models retain or forget factual and relational knowledge when they are continuously edited over time.
RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models
Yanlin Wang, Suiquan Wang, Yanli Wang +4
The paper presents RepoReasoner, a benchmark that evaluates how well large language models can reason about code across multiple files in a repository, testing both fine-grained ex…
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
Chenrui Shi, Yuwei Wu, Yang Liu +5
The paper introduces an Interactive Reward Agent that evaluates GUI task completion by proposing conditions and verifying them using system, application, and GUI tools, and demonst…
CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking
Xiangqun Zhang, Likai Wang, Zekun Qian +2
The paper introduces CD-RMOT-Bench, a benchmark for evaluating how well referring multi-object tracking models trained on one visual domain perform on different, unseen domains, an…
Symbal: Detecting Systematic Misalignments in Model-Generated Captions
Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier +2
The paper presents Symbal, a dual‑stage method that uses off‑the‑shelf foundation models to automatically detect systematic misalignments—recurring caption errors tied to specific…