3 papers
cs.AI2026
SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science
Nithin Somasekharan, Youssef Hassan, Shiyao Lin +5
Large Language Models (LLMs) are increasingly deployed as scientific AI as- sistants, and a growing body of benchmarks evaluates their capabilities across knowledge retrieval, reas…
cs.LG2026
Evaluating Memory Condensation Strategies for Coding Agents in Data-Driven Scientific Discovery
Renuka Chintalapati, Sid Raskar, Anurag Acharya +3
Coding agents accumulate extensive context during long-running tasks, yet fixed context windows force practitioners to choose between truncation and task failure. While numerous me…
cs.CL2026
A Cloud-based Multi-Agentic Workflow for Science
Anurag Acharya, Timothy Vega, Rizwan A. Ashraf +3
As Large Language Models (LLMs) become ubiquitous across various scientific domains, their lack of ability to perform complex tasks like running simulations or to make complex deci…