9 papers
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
Patrick Emami, Sameera Horawalavithana, Truc Nguyen +11
Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI Scientists"…
A Close Look At World Model Recovery In Supervised Fine-Tuned LLM Planners
Patrick Emami, Nan Qiang, Peter Graf
Supervised fine-tuning (SFT) improves end-to-end classical planning in large language models (LLMs), but do these models also learn to represent and reason about the planning probl…
SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science
Nithin Somasekharan, Youssef Hassan, Shiyao Lin +5
Large Language Models (LLMs) are increasingly deployed as scientific AI as- sistants, and a growing body of benchmarks evaluates their capabilities across knowledge retrieval, reas…
Evaluating Memory Condensation Strategies for Coding Agents in Data-Driven Scientific Discovery
Renuka Chintalapati, Sid Raskar, Anurag Acharya +3
Coding agents accumulate extensive context during long-running tasks, yet fixed context windows force practitioners to choose between truncation and task failure. While numerous me…
Born-Qualified: An Autonomous Framework for Deploying Advanced Energy and Electronic Materials
Steven R. Spurgeon, Milad Abolhasani, Frederick Baddour +28
Autonomous science is transforming how we discover materials and chemical systems for advanced energy technologies. However, many initially promising systems never reach deployment…
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics
Nithin Somasekharan, Ling Yue, Yadi Cao +6
Large Language Models (LLMs) have demonstrated strong performance across general NLP tasks, but their utility in automating numerical experiments of complex physical system -- a cr…