◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Da Yu

5 papers hereh-index 478 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author4

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG2
  • cs.CR1
same name
  • Da Yu — 3 papers, h 3
  • Da Yu — 2 papers, h 13
  • Da Yu — 1 paper, h 2
  • Da Yu — 1 paper, h 3

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

5 papers

cs.CL2026

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aaron Blakeman +571

We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 t…

cs.CL2025

Scaling Embedding Layers in Language Models

Da Yu, Edith Cohen, Badih Ghazi +5

We propose SCONE (Scalable, Contextualized, Offloaded, N-gram Embedding), a new method for extending input embedding layers to enhance language model performance. To av…

cs.CR2025

VaultGemma: A Differentially Private Gemma Model

Amer Sinha, Thomas Mesnard, Ryan McKenna +18

We introduce VaultGemma 1B, a 1 billion parameter model within the Gemma family, fully trained with differential privacy. Pretrained on the identical data mixture used for the Gemm…

cs.LG2025

Urania: Differentially Private Insights into AI Use

Daogao Liu, Edith Cohen, Badih Ghazi +8

We introduce Urania, a novel framework for generating insights about LLM chatbot interactions with rigorous differential privacy (DP) guarantees. The framework employs a private…

cs.LG2025

Scaling Laws for Differentially Private Language Models

Ryan McKenna, Yangsibo Huang, Amer Sinha +9

Scaling laws have emerged as important components of large language model (LLM) training as they can predict performance gains through scale, and provide guidance on important hype…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.