11 papers
Operadic consistency: a label-free signal for compositional reasoning failures in LLMs
Nathaniel Bottman, Yinhong Liu, Kyle Richardson
Detecting LLM reasoning failures at inference time without ground-truth labels has motivated a wide range of confidence baselines, including self-consistency, semantic entropy, and…
Operads for compositional reasoning in LLMs
Nathaniel Bottman, Kyle Richardson
Question decomposition, i.e. breaking a complex query into simpler sub-queries whose answers are composed to produce a final answer, is a widely used strategy for improving LLM rea…
ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery
Haofei Yu, Jiaxuan You, Peter Clark +2
Scientific artifacts such as models and datasets are foundations for research. With the rapid growth of platforms like HuggingFace, researchers now have access to a large number of…
Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis
Junyan Cheng, Kyle Richardson, Peter Chin
Large language model (LLM) agents are increasingly tasked with complex real-world analysis (e.g., in financial forecasting, scientific discovery), yet their reasoning suffers from…
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
Jonathan Bragg, Mike D'Arcy, Nishant Balepur +36
AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions o…
TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents
Haofei Yu, Keyang Xuan, Fenghai Li +6
Automatic research with Large Language Models (LLMs) is rapidly gaining importance, driving the development of increasingly complex workflows involving multi-agent systems, plannin…