works on

From the 1 of 27 linked papers with an AI index.

collaborators

27 papers

cs.AI2026

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

Taolin Han, Yuchen Zhang, Jinghang Wang +22

Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce S…

cs.AI2026

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

Xuan Ren, Weiqi Zhai, Tianle Pu +3

Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasonin…

cs.MA2026

Living-Harness Is an Interactive-Agent Evolver

Yuetian Du, Yucheng Wang, He Xu +9

The paper introduces Living-Harness, a self‑evolving harness for large language model agents that updates procedural knowledge from episode feedback, enabling the agent to avoid re…

cs.CL2026

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

Xinke Tong, Xuanming Zhang, Tianyi Tang +10

Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and mult…

cs.CL2026

Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery

Tianyun Zhong, Wangyi Jiang, Wei Wang +15

Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured.…

cs.RO2026

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

Tao Lin, Yuxin Du, Yiran Mao +13

Vision-Language-Action (VLA) models are commonly pretrained on robot demonstrations by jointly mapping visual observations and language instructions to actions. However, dense visu…