5 papers · 1 filter
On Stable Long-Form Generation: Benchmarking and Mitigating Length Volatility
Zhitao He, Haolin Yang, Rui Min +2
Large Language Models (LLMs) excel at long-context understanding but exhibit significant limitations in long-form generation. Existing studies primarily focus on single-generation…
Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight
Haolin Yang, Hakaze Cho, Kaize Ding +1
Large Language Models (LLMs) can perform new tasks from in-context demonstrations, a phenomenon known as in-context learning (ICL). Recent work suggests that these demonstrations a…
Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
Haolin Yang, Hakaze Cho, Naoya Inoue
We investigate the mechanistic underpinnings of in-context learning (ICL) in large language models by reconciling two dominant perspectives: the component-level analysis of attenti…
Neuro-Symbolic Synergy for Interactive World Modeling
Hongyu Zhao, Siyu Zhou, Haolin Yang +2
Large language models (LLMs) exhibit strong general-purpose reasoning capabilities, yet they frequently hallucinate when used as world models (WMs), where strict compliance with de…
Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
Haolin Yang, Hakaze Cho, Yiqiao Zhong +1
The unusual properties of in-context learning (ICL) have prompted investigations into the internal mechanisms of large language models. Prior work typically focuses on either speci…