9 citations · 9 across the 4 of their papers we have counts for
4 papers
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective
Jiayu Liu, Zhenya Huang, Wei Dai +7
Although large language models (LLMs) show promise in solving complex mathematical tasks, existing evaluation paradigms rely solely on a coarse measure of overall answer accuracy,…
MMATH: A Multilingual Benchmark for Mathematical Reasoning
Wenyang Luo, Wayne Xin Zhao, Jing Sha +2
The advent of large reasoning models, such as OpenAI o1 and DeepSeek R1, has significantly advanced complex reasoning tasks. However, their capabilities in multilingual complex rea…
Bit-mask Robust Contrastive Knowledge Distillation for Unsupervised Semantic Hashing
Liyang He, Zhenya Huang, Jiayu Liu +4
Unsupervised semantic hashing has emerged as an indispensable technique for fast image search, which aims to convert images into binary hash codes without relying on labels. Recent…
Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning
Beichen Zhang, Kun Zhou, Xilin Wei +4
Chain-of-thought prompting~(CoT) and tool augmentation have been validated in recent work as effective practices for improving large language models~(LLMs) to perform step-by-step…