7 papers
Multi-Agent Transactive Memory
To Eun Kim, Xuhong He, Dishank Jain +3
The decentralized deployment of LLM agents with diverse capabilities across diverse tasks motivates infrastructure for knowledge sharing across heterogeneous agent populations. Jus…
Offline Preference-Based Trajectory Evaluation
Fernando Diaz
Offline evaluation of agentic systems often collapses trajectories to terminal success, discarding information about partial progress and inducing widespread ties, creating substan…
Characterizing Cultural Localization in AI-Generated Stories
Shaily Bhatt, Supriti Vijay, Jeremiah Milbauer +1
The global use of artificial intelligence has increased interest in assessing the ability to generate culturally localized content, including stories. Cultural localization in stor…
Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation
Xuhong He, To Eun Kim, Maik Fröbe +3
Tip-of-the-Tongue (ToT) retrieval benchmarks have largely focused on English, limiting their applicability to multilingual information access. In this work, we construct multilingu…
Evaluation of Agents under Simulated AI Marketplace Dynamics
To Eun Kim, Alireza Salemi, Hamed Zamani +1
Modern information access ecosystems consist of mixtures of systems, such as retrieval systems and large language models, and increasingly rely on marketplaces to mediate access to…
Overview of the TREC 2025 Tip-of-the-Tongue track
Jaime Arguello, Fernando Diaz, Maik Fröebe +2
Tip-of-the-tongue (ToT) known-item retrieval involves re-finding an item for which the searcher does not reliably recall an identifier. ToT information requests (or queries) are ve…