◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Matilda Orona

2 papers hereh-index 00 citations2 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2

Across the 2 of 2 papers where every author was matched, so the position is known.

fields
  • cs.AI2

identity via Semantic Scholar / OpenAlex

works on
ai agents 1evaluation methodology 1failure analysis 1research automation 1shadow evaluation 1

From the 1 of 2 linked papers with an AI index.

collaborators

2 papers

cs.AI2026

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21

The paper evaluates whether current AI agents can independently conduct open‑ended AI research by having them attempt to solve the central questions of two unpublished NeurIPS subm…

cs.AI2026

Life After Benchmark Saturation: A Case Study of CORE-Bench

Nitya Nadgir, Sayash Kapoor, Kangheng Liu +11

When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privileges accuracy and misses the opportunity t…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.