29 citations · 29 across the 24 of their papers we have counts for
5 papers · 1 filter
GeoDecider: An Evidence-Grounded Agent for Geological Interpretation via Deliberative Reasoning
Jiahao Wang, Mingyue Cheng, Yitong Zhou +6
Geological interpretation infers subsurface properties and structures from indirect geophysical observations. Well-log classification provides a measurable setting by assigning geo…
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
Daoyu Wang, Mingyue Cheng, Shuo Yu +4
Understanding and reasoning on the large-scale scientific literature is a crucial touchstone for large language model (LLM) based agents. However, existing works are mainly restric…
TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning
Chuang Jiang, Mingyue Cheng, Xiaoyu Tao +3
Table reasoning requires models to jointly perform comprehensive semantic understanding and precise numerical operations. Although recent large language model (LLM)-based methods h…
TestAgent: An Adaptive and Intelligent Expert for Human Assessment
Junhao Yu, Yan Zhuang, YuXuan Sun +5
Accurately assessing internal human states is key to understanding preferences, offering personalized services, and identifying challenges in real-world applications. Originating f…
am-ELO: A Stable Framework for Arena-based LLM Evaluation
Zirui Liu, Jiatong Li, Yan Zhuang +5
Arena-based evaluation is a fundamental yet significant evaluation paradigm for modern AI models, especially large language models (LLMs). Existing framework based on ELO rating sy…