agent evaluation 1benchmark consolidation 1capability scaling 1cross-benchmark analysis 1dataset infrastructure 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Stefan Krsteski, Charlotte Meyer, Guillaume Allegre +2
The paper introduces Messier, a unified corpus of 957,253 standardized evaluation records spanning thousands of agents, tasks, and benchmarks, to enable cross‑benchmark analysis an…
cs.SE2025
MMORE: Massive Multimodal Open RAG & Extraction
Alexandre Sallinen, Stefan Krsteski, Paul Teiletche +7
We introduce MMORE, an open-source pipeline for Massive Multimodal Open RetrievalAugmented Generation and Extraction, designed to ingest, transform, and retrieve knowledge from het…
cs.CL2025
Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems
Shang-Chi Tsai, Yun-Nung Chen
With the advancement of large language models, many dialogue systems are now capable of providing reasonable and informative responses to patients' medical conditions. However, whe…