most citedIs Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CL20261 cited

Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers

Kaiyu He, Zhang Mian, Peilin Wu +2

While Large Language Models (LLMs) excel at factual retrieval, they often struggle with the "curse of two-hop reasoning" in compositional tasks. Recent research suggests that param…

cs.CL2025

GEAR: A General Evaluation Framework for Abductive Reasoning

Kaiyu He, Peilin Wu, Mian Zhang +4

Since the advent of large language models (LLMs), research has focused on instruction following and deductive reasoning. A central question remains: can these models discover new k…

cs.SE2025

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Shuo Yan, Ruochen Li, Ziming Luo +11

Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery. However, their capability in the fundamental yet crucial task of reprodu…

cs.CL2025

Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty

Peilin Wu, Mian Zhang, Xinlu Zhang +2

Agentic Retrieval-Augmented Generation (RAG) systems enhance Large Language Models (LLMs) by enabling dynamic, multi-step reasoning and information retrieval. However, these system…

cs.CL2025

Do Retrieval-Augmented Language Models Adapt to Varying User Needs?

Peilin Wu, Xinlu Zhang, Wenhao Yu +3

Recent advancements in Retrieval-Augmented Language Models (RALMs) have demonstrated their efficacy in knowledge-intensive tasks. However, existing evaluation benchmarks often assu…