activity
20242026
collaborators
Showing cs.AIShow all

7 papers · 1 filter

cs.AI2026

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

Zijian Wang, Hanqi Li, Ziyue Yang +17

AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside…

cs.AI2026

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

Wanhao Liu, Jiaqing Xie, Qian Tan +10

As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdisciplinary, multimodal, and…

cs.AI2026

DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation

Ziyue Yang, Da Ma, Hanqi Li +8

As scientific literature grows rapidly, automated survey generation has become a key capability for AI scientists and human researchers. However, existing systems suffer from limit…

cs.AI2026

Diagnosing CFG Interpretation in LLMs

Hanqi Li, Lu Chen, Kai Yu

As LLMs are increasingly integrated into agentic systems, they must adhere to dynamically defined, machine-interpretable interfaces. We evaluate LLMs as in-context interpreters: gi…

cs.AI2026

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding

Situo Zhang, Yifan Zhang, Zichen Zhu +6

Charts are ubiquitous in scientific and financial literature for presenting structured data. However, chart reasoning remains challenging for multimodal large language models (MLLM…

cs.AI2025

ProgRM: Build Better GUI Agents with Progress Rewards

Danyang Zhang, Situo Zhang, Ziyue Yang +5

LLM-based (Large Language Model) GUI (Graphical User Interface) agents can potentially reshape our daily lives significantly. However, current LLM-based GUI agents suffer from the…