8 papers
Living-Harness Is an Interactive-Agent Evolver
Yuetian Du, Yucheng Wang, He Xu +9
The paper introduces Living-Harness, a self‑evolving harness for large language model agents that updates procedural knowledge from episode feedback, enabling the agent to avoid re…
Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
Xinke Tong, Xuanming Zhang, Tianyi Tang +10
Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and mult…
Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery
Tianyun Zhong, Wangyi Jiang, Wei Wang +15
Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured.…
Logics-Parsing-Omni Technical Report
Xin An, Jingyi Cai, Xiangyang Chen +22
Addressing the challenges of fragmented task definitions and the heterogeneity of unstructured data in multimodal parsing, this paper proposes the Omni Parsing framework. This fram…
Linking Perception, Confidence and Accuracy in MLLMs
Yuetian Du, Yucheng Wang, Rongyu Zhang +5
Recent advances in Multi-modal Large Language Models (MLLMs) have predominantly focused on enhancing visual perception to improve accuracy. However, a critical question remains une…
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
Xiang Zheng, Weiqi Zhai, Wei Wang +15
Recent large language models (LLMs) achieve near-saturation accuracy on many established mathematical reasoning benchmarks, raising concerns about their ability to diagnose genuine…