most citedCollapse of Self-trained Language Models

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL2024

Rethinking Thinking Tokens: Understanding Why They Underperform in Practice

Sreeram Vennam, David Valente, David Herel +1

Thinking Tokens (TT) have been proposed as an unsupervised method to facilitate reasoning in language models. However, despite their conceptual appeal, our findings show that TTs m…

cs.CL2024

Time Awareness in Large Language Models: Benchmarking Fact Recall Across Time

David Herel, Vojtech Bartek, Jiri Jirak +1

Who is the US President? The answer changes depending on when the question is asked. While large language models (LLMs) are evaluated on various reasoning tasks, they often miss a…

cs.CL2024

Thinking Tokens for Language Modeling

David Herel, Tomas Mikolov

How much is 56 times 37? Language models often make mistakes in these types of difficult calculations. This is usually explained by their inability to perform complex reasoning. Si…

cs.CL20241 cited

Collapse of Self-trained Language Models

David Herel, Tomas Mikolov

In various fields of knowledge creation, including science, new ideas often build on pre-existing information. In this work, we explore this concept within the context of language…

cs.LG2024

On Difficulties of Attention Factorization through Shared Memory

Uladzislau Yorsh, Martin Holeňa, Ondřej Bojar +1

Transformers have revolutionized deep learning in numerous fields, including natural language processing, computer vision, and audio processing. Their strength lies in their attent…

cs.CL2023

Advancing State of the Art in Language Modeling

David Herel, Tomas Mikolov

Generalization is arguably the most important goal of statistical language modeling research. Publicly available benchmarks and papers published with an open-source code have been…