1 citations · 1 across the 3 of their papers we have counts for
6 papers
Rethinking Thinking Tokens: Understanding Why They Underperform in Practice
Sreeram Vennam, David Valente, David Herel +1
Thinking Tokens (TT) have been proposed as an unsupervised method to facilitate reasoning in language models. However, despite their conceptual appeal, our findings show that TTs m…
Time Awareness in Large Language Models: Benchmarking Fact Recall Across Time
David Herel, Vojtech Bartek, Jiri Jirak +1
Who is the US President? The answer changes depending on when the question is asked. While large language models (LLMs) are evaluated on various reasoning tasks, they often miss a…
Thinking Tokens for Language Modeling
David Herel, Tomas Mikolov
How much is 56 times 37? Language models often make mistakes in these types of difficult calculations. This is usually explained by their inability to perform complex reasoning. Si…
Collapse of Self-trained Language Models
David Herel, Tomas Mikolov
In various fields of knowledge creation, including science, new ideas often build on pre-existing information. In this work, we explore this concept within the context of language…
On Difficulties of Attention Factorization through Shared Memory
Uladzislau Yorsh, Martin Holeňa, Ondřej Bojar +1
Transformers have revolutionized deep learning in numerous fields, including natural language processing, computer vision, and audio processing. Their strength lies in their attent…
Advancing State of the Art in Language Modeling
David Herel, Tomas Mikolov
Generalization is arguably the most important goal of statistical language modeling research. Publicly available benchmarks and papers published with an open-source code have been…