3 papers
cs.CL2024
Dynamic layer selection in decoder-only transformers
Theodore Glavas, Joud Chataoui, Florence Regol +4
The vast size of Large Language Models (LLMs) has prompted a search to optimize inference. One effective approach is dynamic inference, which adapts the architecture to the sample-…
cs.CV2024
Unsupervised Domain Adaptation Approaches for Chessboard Recognition
Wassim Jabbour, Enzo Benoit-Jeannin, Oscar Bedford +1
Chess involves extensive study and requires players to keep manual records of their matches, a process which is time-consuming and distracting. The lack of high-quality labeled pho…
cs.CL2024
Scavenging Hyena: Distilling Transformers into Long Convolution Models
Tokiniaina Raharison Ralambomihanta, Shahrad Mohammadzadeh, Mohammad Sami Nur Islam +2
The rapid evolution of Large Language Models (LLMs), epitomized by architectures like GPT-4, has reshaped the landscape of natural language processing. This paper introduces a pion…