8 papers
scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
Ian Diks, Zhen Yang, Arjun Banerjee +2
Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxili…
SPARC: A Multi-Agent System for Electrical Circuit Question Answering
Mushtari Sadia, Zhenning Yang, Umme Habiba Lamia +3
Electrical circuit diagram QA tasks require complex mathematical reasoning, which remains challenging for multimodal LLMs. We present SPARC, a multi-agent system that answers quest…
Ambig-IaC: Multi-level Disambiguation for Interactive Cloud Infrastructure-as-Code Synthesis
Zhenning Yang, Kaden Gruizenga, Tongyuan Miao +3
The scale and complexity of modern cloud infrastructure have made Infrastructure-as-Code (IaC) essential for managing deployments. While large Language models (LLMs) are increasing…
SQUiD: Synthesizing Relational Databases from Unstructured Text
Mushtari Sadia, Zhenning Yang, Yunming Xiao +2
Relational databases are central to modern data management, yet most data exists in unstructured forms like text documents. To bridge this gap, we leverage large language models (L…
Automated Cloud Infrastructure-as-Code Reconciliation with AI Agents
Zhenning Yang, Hui Guan, Victor Nicolet +4
Cloud infrastructure is managed through a mix of interfaces -- traditionally, cloud consoles, command-line interfaces (CLI), and SDKs are the tools of choice. Recently, Infrastruct…
Cloud Infrastructure Management in the Age of AI Agents
Zhenning Yang, Archit Bhatnagar, Yiming Qiu +6
Cloud infrastructure is the cornerstone of the modern IT industry. However, managing this infrastructure effectively requires considerable manual effort from the DevOps engineering…