activity
20242026
most citedAstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2026

Operadic consistency: a label-free signal for compositional reasoning failures in LLMs

Nathaniel Bottman, Yinhong Liu, Kyle Richardson

Detecting LLM reasoning failures at inference time without ground-truth labels has motivated a wide range of confidence baselines, including self-consistency, semantic entropy, and…

cs.CL2026

Operads for compositional reasoning in LLMs

Nathaniel Bottman, Kyle Richardson

Question decomposition, i.e. breaking a complex query into simpler sub-queries whose answers are composed to produce a final answer, is a widely used strategy for improving LLM rea…

cs.CL2025

TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents

Haofei Yu, Keyang Xuan, Fenghai Li +6

Automatic research with Large Language Models (LLMs) is rapidly gaining importance, driving the development of increasingly complex workflows involving multi-agent systems, plannin…

cs.CL2025

Event Causality Identification with Synthetic Control

Haoyu Wang, Fengze Liu, Jiayao Zhang +2

Event causality identification (ECI), a process that extracts causal relations between events from text, is crucial for distinguishing causation from correlation. Traditional appro…

cs.CL2025

Understanding the Logic of Direct Preference Alignment through Logic

Kyle Richardson, Vivek Srikumar, Ashish Sabharwal

Recent direct preference alignment algorithms (DPA), such as DPO, have shown great promise in aligning large language models to human preferences. While this has motivated the deve…

cs.CL2024

Paloma: A Benchmark for Evaluating Language Model Fit

Ian Magnusson, Akshita Bhagia, Valentin Hofmann +13

Evaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains--varying distr…