4 papers
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
Atsuyuki Miyai, Mashiro Toyooka, Zaiying Zhao +3
This paper introduces the first systematic evaluation framework for quantifying the quality and risks of papers written by modern coding agents. While AI-driven paper writing has b…
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
Atsuyuki Miyai, Mashiro Toyooka, Takashi Otonari +2
Understanding the current capabilities and risks of AI Scientist systems (autoresearch) is essential for ensuring trustworthy and sustainable AI-driven scientific progress while pr…
A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
Mashiro Toyooka, Kiyoharu Aizawa, Yoko Yamakata
Large Language Models (LLMs) are trained on a vast amount of procedural texts, but they do not directly observe real-world phenomena. In the context of cooking recipes, this poses…
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
Atsuyuki Miyai, Zaiying Zhao, Kazuki Egashira +9
Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent and general framework for automating w…