◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Tri Dao

11 papers hereh-index 99.7k citations15 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author5
  • last author5

Across the 11 of 11 papers where every author was matched, so the position is known.

fields
  • cs.LG7
  • cs.AI1
  • cs.CL1
  • cs.RO1
  • q-bio.GN1
same name
  • Tri Dao — 29 papers, h 27
  • Tri Dao — 9 papers, h 3
  • Tri Dao — 8 papers, h 7
  • Tri Dao — 6 papers, h 4
  • Tri Dao — 3 papers, h 4

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20232026
most citedMamba: Linear-Time Sequence Modeling with Selective State Spaces

1.1k citations · 1.2k across the 10 of their papers we have counts for

collaborators
Showing 2024 · cs.LGShow all

4 papers · 2 filters

cs.LG2024★ 2 cited

The Mamba in the Llama: Distilling and Accelerating Hybrid Models

Junxiong Wang, Daniele Paliotta, Avner May +2

Linear RNN architectures, like Mamba, can be competitive with Transformer models in language modeling while having advantageous deployment characteristics. Given the focus on train…

cs.LG2024

Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers

Sukjun Hwang, Aakash Lahoti, Tri Dao +1

A wide array of sequence models are built on a framework modeled after Transformers, comprising alternating sequence mixer and channel mixer layers. This paper studies a unifying m…

cs.LG2024★ 9 cited

An Empirical Study of Mamba-based Language Models

Roger Waleffe, Wonmin Byeon, Duncan Riach +13

Selective state-space models (SSMs) like Mamba overcome some of the shortcomings of Transformers, such as quadratic computational complexity with sequence length and large inferenc…

cs.LG2024★ 78 cited

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Tri Dao, Albert Gu

While Transformers have been the main architecture behind deep learning's success in language modeling, state-space models (SSMs) such as Mamba have recently been shown to match or…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.