most citedIs Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL20261 cited

Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers

Kaiyu He, Zhang Mian, Peilin Wu +2

While Large Language Models (LLMs) excel at factual retrieval, they often struggle with the "curse of two-hop reasoning" in compositional tasks. Recent research suggests that param…

cs.CL2025

GEAR: A General Evaluation Framework for Abductive Reasoning

Kaiyu He, Peilin Wu, Mian Zhang +4

Since the advent of large language models (LLMs), research has focused on instruction following and deductive reasoning. A central question remains: can these models discover new k…

cs.SE2025

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Shuo Yan, Ruochen Li, Ziming Luo +11

Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery. However, their capability in the fundamental yet crucial task of reprodu…

cs.CL2025

Semantic Pivots Enable Cross-Lingual Transfer in Large Language Models

Kaiyu He, Tong Zhou, Yubo Chen +4

Large language models (LLMs) demonstrate remarkable ability in cross-lingual tasks. Understanding how LLMs acquire this ability is crucial for their interpretability. To quantify t…

cs.CL2025

From Reasoning to Learning: A Survey on Hypothesis Discovery and Rule Learning with Large Language Models

Kaiyu He, Zhiyu Chen

Since the advent of Large Language Models (LLMs), efforts have largely focused on improving their instruction-following and deductive reasoning abilities, leaving open the question…