#evaluation metrics
31 papers · 1 filter
When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence
Zongheng Guo, Tao Chen, Tianli Li +6
The paper defines and quantifies derived-feature over‑trust (DFOT) where large language models treat derived measurements as direct facts, using physiological sensing (PPG vs. ECG)…
Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3
Jens Lehmann, Andrei Aioanei, Sahar Vahdati
The paper presents Tycho, a coding-agent that actively builds and uses programmatic world models to infer game rules and improve action efficiency in the ARC-AGI-3 benchmark, achie…
MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware
Xianpeng Zhang, Jiahua Yang, Dongyu Chen +7
The paper presents a benchmark for multimodal long-document summarization and a two-stage training framework (MMLDSum-LLM) that incorporates visual-alignment and keyword-aware loss…
Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents
Jiwon Jang, Kisu Yang, Heuiseok Lim +1
The paper evaluates 4-bit post‑training quantization of multi‑turn, tool‑calling LLM agents and finds that while standard scores remain unchanged, quantization substantially increa…
APEX-Accounting
Julien Benchek, Austin Bennett, Jasmin Kern +8
The paper presents APEX-Accounting, a benchmark for evaluating how well advanced language models can perform real accounting tasks such as reconciliation, expense accrual, transact…
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
Weijie Wu, Junbo Li, Lin Li +2
The paper introduces MMAC, a large benchmark of 5,638 audio clips designed to evaluate audio captioning models across multiple capability categories and evaluation dimensions, focu…