6 papers
TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate
Erica Zhang, Fangzhao Zhang, Aneesh Pappu +5
Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation. It is also a canonical testbed for agentic languag…
Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
William Lehn-Schiøler, Magnus Ruud Kjær, Rahul Thapa +10
EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust. We apply To…
Unlocking LLM Creativity in Science through Analogical Reasoning
Andrew Shen, Shaul Druckmann, James Zou
Autonomous science promises to augment scientific discovery, particularly in complex fields like biomedicine. However, this requires AI systems that can consistently generate novel…
A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
Wei-Chieh Huang, Weizhi Zhang, Yueqing Liang +57
Research in artificial intelligence is shifting from model innovations and benchmark scores towards problem definition and rigorous real-world evaluation. As the field enters the "…
Reliable and Responsible Foundation Models: A Comprehensive Survey
Xinyu Yang, Junlin Han, Rishi Bommasani +49
Foundation models, including Large Language Models (LLMs), Multimodal Large Language Models (MLLMs), Image Generative Models (i.e, Text-to-Image Models and Image-Editing Models), a…
A Theoretical Framework for Prompt Engineering: Approximating Smooth Functions with Transformer Prompts
Ryumei Nakada, Wenlong Ji, Tianxi Cai +2
Prompt engineering has emerged as a powerful technique for guiding large language models (LLMs) toward desired responses, significantly enhancing their performance across diverse t…