1 citations · 1 across the 3 of their papers we have counts for
5 papers
Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers
Kaiyu He, Zhang Mian, Peilin Wu +2
While Large Language Models (LLMs) excel at factual retrieval, they often struggle with the "curse of two-hop reasoning" in compositional tasks. Recent research suggests that param…
GEAR: A General Evaluation Framework for Abductive Reasoning
Kaiyu He, Peilin Wu, Mian Zhang +4
Since the advent of large language models (LLMs), research has focused on instruction following and deductive reasoning. A central question remains: can these models discover new k…
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
Shuo Yan, Ruochen Li, Ziming Luo +11
Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery. However, their capability in the fundamental yet crucial task of reprodu…
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty
Peilin Wu, Mian Zhang, Xinlu Zhang +2
Agentic Retrieval-Augmented Generation (RAG) systems enhance Large Language Models (LLMs) by enabling dynamic, multi-step reasoning and information retrieval. However, these system…
Do Retrieval-Augmented Language Models Adapt to Varying User Needs?
Peilin Wu, Xinlu Zhang, Wenhao Yu +3
Recent advancements in Retrieval-Augmented Language Models (RALMs) have demonstrated their efficacy in knowledge-intensive tasks. However, existing evaluation benchmarks often assu…