◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Ilyas Chahed

4 papers hereh-index 286 citations4 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author4

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG2

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators

4 papers

cs.LG2026

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training

Jingwei Zuo, Cong Zeng, Ilyas Chahed +6

The training paradigm of large language models has shifted from traditional one-pass training to multi-epoch training, as reasonable reuse of limited high-quality data can improve…

cs.LG2026

Learnable Multipliers: Freeing the Scale of Language Model Matrix Layers

Maksim Velikanov, Ilyas Chahed, Jingwei Zuo +3

Applying weight decay (WD) to matrix layers is standard practice in large-language-model pretraining. Prior work suggests that stochastic gradient noise induces a Brownian-like exp…

cs.CL2025

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance

Jingwei Zuo, Maksim Velikanov, Ilyas Chahed +24

In this report, we introduce Falcon-H1, a new series of large language models (LLMs) featuring hybrid architecture designs optimized for both high performance and efficiency across…

cs.CL2024

Falcon Mamba: The First Competitive Attention-free 7B Language Model

Jingwei Zuo, Maksim Velikanov, Dhia Eddine Rhaiem +4

In this technical report, we present Falcon Mamba 7B, a new base large language model based on the novel Mamba architecture. Falcon Mamba 7B is trained on 5.8 trillion tokens with…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.