3 citations · 5 across the 13 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
ChatBench: From Static Benchmarks to Human-AI Evaluation
Serina Chang, Ashton Anderson, Jake M. Hofman
With the rapid adoption of LLM-based chatbots, there is a pressing need to evaluate what humans and LLMs can achieve together. However, standard benchmarks, such as MMLU, measure L…
cs.CL2023
ICL Markup: Structuring In-Context Learning using Soft-Token Tags
Marc-Etienne Brunet, Ashton Anderson, Richard Zemel
Large pretrained language models (LLMs) can be rapidly adapted to a wide variety of tasks via a text-to-text approach, where the instruction and input are fed to the model in natur…