activity
20182026
most citedLarge Language Models are Zero-Shot Reasoners

1.1k citations · 1.3k across the 91 of their papers we have counts for

collaborators
Showing 2025 · cs.CLShow all

11 papers · 2 filters

cs.CL2025

Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning

Ru Wang, Wei Huang, Qi Cao +3

Test-time reinforcement learning (TTRL) offers a label-free paradigm for adapting models using only synthetic signals at inference, but its success hinges on constructing reliable…

cs.CL2025

Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise

Keno Harada, Lui Yoshida, Takeshi Kojima +2

The performance of Large Language Models (LLMs) is highly sensitive to the prompts they are given. Drawing inspiration from the field of prompt optimization, this study investigate…

cs.CL2025

When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following

Keno Harada, Yudai Yamazaki, Masachika Taniguchi +4

As large language models (LLMs) are increasingly applied to real-world scenarios, it becomes crucial to understand their ability to follow multiple instructions simultaneously. To…

cs.CL2025

Dynamic Injection of Entity Knowledge into Dense Retrievers

Ikuya Yamada, Ryokan Ri, Takeshi Kojima +2

Dense retrievers often struggle with queries involving less-frequent entities due to their limited entity knowledge. We propose the Knowledgeable Passage Retriever (KPR), a BERT-ba…

cs.CL2025

Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence

Gouki Minegishi, Hiroki Furuta, Shohei Taniguchi +2

Transformer-based language models exhibit In-Context Learning (ICL), where predictions are made adaptively based on context. While prior work links induction heads to ICL through a…

cs.CL2025

Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar

Andrew Gambardella, Takeshi Kojima, Yusuke Iwasawa +1

Typical methods for evaluating the performance of language models evaluate their ability to answer questions accurately. These evaluation metrics are acceptable for determining the…