◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Sushant Mehta

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.AI1
  • cs.CL1
  • cs.LG1
same name
  • Sushant Mehta — 1 paper, h 2

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedLatent Multi-Head Attention for Small Language Models

1 citations · 1 across the 3 of their papers we have counts for

collaborators

3 papers

cs.LG2025

Muon: Training and Trade-offs with Latent Attention and MoE

Sushant Mehta, Raj Dandekar, Rajat Dandekar +1

We present a comprehensive theoretical and empirical study of the Muon optimizer for training transformers only with a small to medium decoder (30M - 200M parameters), with an emph…

cs.AI2025

Unifying Mixture of Experts and Multi-Head Latent Attention for Efficient Language Models

Sushant Mehta, Raj Dandekar, Rajat Dandekar +1

We present MoE-MLA-RoPE, a novel architecture combination that combines Mixture of Experts (MoE) with Multi-head Latent Attention (MLA) and Rotary Position Embeddings (RoPE) for ef…

cs.CL2025★ 1 cited

Latent Multi-Head Attention for Small Language Models

Sushant Mehta, Raj Dandekar, Rajat Dandekar +1

We present the first comprehensive study of latent multi-head attention (MLA) for small language models, revealing interesting efficiency-quality trade-offs. Training 30M-parameter…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.