13 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.AI2026
A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline
Kai A. Horstmann, Ethan Lin, Alice A. Robie +2
Agentic AI offers a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that take domain experts days to months…
cs.LG2022★ 13 cited
MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of Behavior
Jennifer J. Sun, Markus Marks, Andrew Ulmer +20
We introduce MABe22, a large-scale, multi-agent video and trajectory benchmark to assess the quality of learned behavior representations. This dataset is collected from a variety o…