5 citations · 5 across the 3 of their papers we have counts for
3 papers
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
Rahul Atul Bhope, Praveen Venkateswaran, K. R. Jayaram +3
Developers using LLMs and LLM-based agents in their applications have provided plenty of anecdotal evidence that in-context-learning (ICL) is fragile. In this paper, we show that i…
Leveraging the Inductive Bias of Large Language Models for Abstract Textual Reasoning
Christopher Michael Rytting, David Wingate
Large natural language models (such as GPT-3 or T5) demonstrate impressive abilities across a range of general NLP tasks. Here, we show that the knowledge embedded in such models p…