activity
20242026
collaborators

11 papers

cs.CL2026

Operadic consistency: a label-free signal for compositional reasoning failures in LLMs

Nathaniel Bottman, Yinhong Liu, Kyle Richardson

Detecting LLM reasoning failures at inference time without ground-truth labels has motivated a wide range of confidence baselines, including self-consistency, semantic entropy, and…

cs.CL2026

Operads for compositional reasoning in LLMs

Nathaniel Bottman, Kyle Richardson

Question decomposition, i.e. breaking a complex query into simpler sub-queries whose answers are composed to produce a final answer, is a widely used strategy for improving LLM rea…

cs.LG2026

ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery

Haofei Yu, Jiaxuan You, Peter Clark +2

Scientific artifacts such as models and datasets are foundations for research. With the rapid growth of platforms like HuggingFace, researchers now have access to a large number of…

cs.AI2026

Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis

Junyan Cheng, Kyle Richardson, Peter Chin

Large language model (LLM) agents are increasingly tasked with complex real-world analysis (e.g., in financial forecasting, scientific discovery), yet their reasoning suffers from…

cs.AI2026

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

Jonathan Bragg, Mike D'Arcy, Nishant Balepur +36

AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions o…

cs.CL2025

TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents

Haofei Yu, Keyang Xuan, Fenghai Li +6

Automatic research with Large Language Models (LLMs) is rapidly gaining importance, driving the development of increasingly complex workflows involving multi-agent systems, plannin…