collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

Yanchao Li, Wanhao Liu, Jiaqing Xie +4

Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach new answers, or produc…

cs.AI2026

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

Wanhao Liu, Jiaqing Xie, Qian Tan +10

As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdisciplinary, multimodal, and…

cs.AI2026

SkillsInjector: Dynamic Skill Context Construction for LLM Agents

Yanchao Li, Wanhao Liu, Ben Gao +5

LLM agents now draw on growing skill libraries to handle complex tasks. However, injecting more skills does not always improve task completion and can even degrade it. Existing met…

cs.AI2025

MolAct: An Agentic RL Framework for Molecular Editing and Property Optimization

Zhuo Yang, Yeyun Chen, Jiaqing Xie +7

Molecular editing and optimization are multi-step problems that require iteratively improving properties while keeping molecules chemically valid and structurally similar. We frame…

cs.AI2025

Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows

Wanghan Xu, Yuhao Zhou, Yifan Zhou +104

Despite advances in scientific AI, a coherent framework for Scientific General Intelligence (SGI)-the ability to autonomously conceive, investigate, and reason across scientific do…