works on

From the 2 of 34 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

Anqi Zou, Han Deng, Chengyu Zhang +9

Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over co…

cs.AI2026

MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs

Xiangyu Zhao, Wanghan Xu, Bo Liu +7

The rapid advancement of multimodal large language models (MLLMs) offers new opportunities for complex scientific challenges, yet their application in earth science-especially at t…

cs.AI2026

Owl-AuraID 1.0: An Intelligent System for Autonomous Scientific Instrumentation and Scientific Data Analysis

Han Deng, Anqi Zou, Hanling Zhang +14

Scientific discovery increasingly depends on high-throughput characterization, yet automation is hindered by proprietary GUIs and the limited generalizability of existing API-based…

cs.AI2026

Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System

Wanghan Xu, Wenlong Zhang, Fenghua Ling +8

Meta-analysis is a systematic research methodology that synthesizes data from multiple existing studies to derive comprehensive conclusions. This approach not only mitigates limita…

cs.AI2025

Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows

Wanghan Xu, Yuhao Zhou, Yifan Zhou +104

Despite advances in scientific AI, a coherent framework for Scientific General Intelligence (SGI)-the ability to autonomously conceive, investigate, and reason across scientific do…

cs.AI2025

Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning

Yuhao Zhou, Yiheng Wang, Xuming He +26

Scientific discoveries increasingly rely on complex multimodal reasoning based on information-intensive scientific data and domain-specific expertise. Empowered by expert-level sci…