7 papers · 1 filter
Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness
Zijian Wang, Hanqi Li, Ziyue Yang +17
AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside…
OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields
Wanhao Liu, Jiaqing Xie, Qian Tan +10
As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdisciplinary, multimodal, and…
DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation
Ziyue Yang, Da Ma, Hanqi Li +8
As scientific literature grows rapidly, automated survey generation has become a key capability for AI scientists and human researchers. However, existing systems suffer from limit…
Diagnosing CFG Interpretation in LLMs
Hanqi Li, Lu Chen, Kai Yu
As LLMs are increasingly integrated into agentic systems, they must adhere to dynamically defined, machine-interpretable interfaces. We evaluate LLMs as in-context interpreters: gi…
CharTool: Tool-Integrated Visual Reasoning for Chart Understanding
Situo Zhang, Yifan Zhang, Zichen Zhu +6
Charts are ubiquitous in scientific and financial literature for presenting structured data. However, chart reasoning remains challenging for multimodal large language models (MLLM…
ProgRM: Build Better GUI Agents with Progress Rewards
Danyang Zhang, Situo Zhang, Ziyue Yang +5
LLM-based (Large Language Model) GUI (Graphical User Interface) agents can potentially reshape our daily lives significantly. However, current LLM-based GUI agents suffer from the…