#evaluation metrics

topicevaluation metrics

31 papers · 1 filter

cs.AI2026

When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence

Zongheng Guo, Tao Chen, Tianli Li +6

The paper defines and quantifies derived-feature over‑trust (DFOT) where large language models treat derived measurements as direct facts, using physiological sensing (PPG vs. ECG)…

cs.AI2026

Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3

Jens Lehmann, Andrei Aioanei, Sahar Vahdati

The paper presents Tycho, a coding-agent that actively builds and uses programmatic world models to infer game rules and improve action efficiency in the ARC-AGI-3 benchmark, achie…

cs.AI2026

MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware

Xianpeng Zhang, Jiahua Yang, Dongyu Chen +7

The paper presents a benchmark for multimodal long-document summarization and a two-stage training framework (MMLDSum-LLM) that incorporates visual-alignment and keyword-aware loss…

cs.LG2026

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

Jiwon Jang, Kisu Yang, Heuiseok Lim +1

The paper evaluates 4-bit post‑training quantization of multi‑turn, tool‑calling LLM agents and finds that while standard scores remain unchanged, quantization substantially increa…

cs.CL2026

APEX-Accounting

Julien Benchek, Austin Bennett, Jasmin Kern +8

The paper presents APEX-Accounting, a benchmark for evaluating how well advanced language models can perform real accounting tasks such as reconciliation, expense accrual, transact…

cs.SD2026

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

Weijie Wu, Junbo Li, Lin Li +2

The paper introduces MMAC, a large benchmark of 5,638 audio clips designed to evaluate audio captioning models across multiple capability categories and evaluation dimensions, focu…