agent evaluation 1benchmark consolidation 1capability scaling 1cross-benchmark analysis 1dataset infrastructure 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Stefan Krsteski, Charlotte Meyer, Guillaume Allegre +2
The paper introduces Messier, a unified corpus of 957,253 standardized evaluation records spanning thousands of agents, tasks, and benchmarks, to enable cross‑benchmark analysis an…
cs.LG2025
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing
Minghui Liu, Tahseen Rabbani, Tony O'Halloran +5
Transformer-based large language models (LLMs) use the key-value (KV) cache to significantly accelerate inference by storing the key and value embeddings of past tokens. However, t…