collaborators

29 papers

cs.AI2026

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

Xuan Ren, Weiqi Zhai, Tianle Pu +3

Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasonin…

cs.MA2026

Living-Harness Is an Interactive-Agent Evolver

Yuetian Du, Yucheng Wang, He Xu +9

The paper introduces Living-Harness, a self‑evolving harness for large language model agents that updates procedural knowledge from episode feedback, enabling the agent to avoid re…

cs.CL2026

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

Xinke Tong, Xuanming Zhang, Tianyi Tang +10

Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and mult…

cs.CL2026

Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery

Tianyun Zhong, Wangyi Jiang, Wei Wang +15

Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured.…

cs.LG2026

Discovering Millions of Interpretable Features with Sparse Autoencoders

XinYang He, Wei Wang, Bing Zhao +5

Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs…

cs.CV2026

Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation

Niantong Li, Guangzheng Hu, Weixu Qiao +35

Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no…