1 citations · 1 across the 4 of their papers we have counts for
4 papers
Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers
Kaiyu He, Zhang Mian, Peilin Wu +2
While Large Language Models (LLMs) excel at factual retrieval, they often struggle with the "curse of two-hop reasoning" in compositional tasks. Recent research suggests that param…
ELPO: Ensemble Learning Based Prompt Optimization for Large Language Models
Qing Zhang, Bing Xu, Xudong Zhang +9
The remarkable performance of Large Language Models (LLMs) highly relies on crafted prompts. However, manual prompt engineering is a laborious process, creating a core bottleneck f…
GEAR: A General Evaluation Framework for Abductive Reasoning
Kaiyu He, Peilin Wu, Mian Zhang +4
Since the advent of large language models (LLMs), research has focused on instruction following and deductive reasoning. A central question remains: can these models discover new k…
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
Shuo Yan, Ruochen Li, Ziming Luo +11
Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery. However, their capability in the fundamental yet crucial task of reprodu…