2 papers
cs.AI2026
Recovering Wasted Compute in Autoresearch Agents
Au Kwok Chun, Abhigyan Acherjee, Amrutha Rao +4
A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. Such agents have inspired large industry invest…
cs.LG2026
Toward Human Rights Benchmarking for LLMs: A Pilot Methodology
Savannah Thais, Wm. Matthew Kennedy, Abhigyan Acherjee +3
Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can…