4 citations · 4 across the 2 of their papers we have counts for
6 papers
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
Daoyu Wang, Mingyue Cheng, Shuo Yu +4
Understanding and reasoning on the large-scale scientific literature is a crucial touchstone for large language model (LLM) based agents. However, existing works are mainly restric…
SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment
Yuqing Huang, Rongyang Zhang, Qimeng Wang +9
Recent advancements in large language models (LLMs) have revolutionized natural language processing through their remarkable capabilities in understanding and executing diverse tas…
Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables
Yitong Zhou, Mingyue Cheng, Qingyang Mao +7
With the widespread application of multimodal large language models in scientific intelligence, there is an urgent need for more challenging evaluation benchmarks to assess their a…
Enhancing Table Recognition with Vision LLMs: A Benchmark and Neighbor-Guided Toolchain Reasoner
Yitong Zhou, Mingyue Cheng, Qingyang Mao +2
Pre-trained foundation models have recently made significant progress in table-related tasks such as table understanding and reasoning. However, recognizing the structure and conte…
TableTime: Reformulating Time Series Classification as Training-Free Table Understanding with Large Language Models
Jiahao Wang, Mingyue Cheng, Qingyang Mao +5
Large language models (LLMs) have demonstrated their effectiveness in multivariate time series classification (MTSC). Effective adaptation of LLMs for MTSC necessitates informative…
ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models
Yuqing Huang, Rongyang Zhang, Xuesong He +15
There is a growing interest in the role that LLMs play in chemistry which lead to an increased focus on the development of LLMs benchmarks tailored to chemical domains to assess th…