◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Michael Y Hu

New York University

4 papers hereh-index 10278 citations14 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author2

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL3
  • cs.LG1
affiliations
  • New York University
Homepage

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.CL2026

Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time

Michael Y. Hu, Apurva Gandhi, Kyunghyun Cho +2

Data mixing decides how to combine different sources or types of data and is a consequential problem throughout language model training. In pretraining, data composition is a key d…

cs.CL2025

Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check

Nicholas Lourie, Michael Y. Hu, Kyunghyun Cho

Downstream scaling laws aim to predict task performance at larger scales from the model's performance at smaller scales. Whether such prediction should be possible is unclear: some…

cs.CL2025

Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases

Michael Y. Hu, Jackson Petty, Chuan Shi +2

Pretraining language models on formal language can improve their acquisition of natural language. Which features of the formal language impart an inductive bias that leads to effec…

cs.LG2025

Aioli: A Unified Optimization Framework for Language Model Data Mixing

Mayee F. Chen, Michael Y. Hu, Nicholas Lourie +2

Language model performance depends on identifying the optimal mixture of data groups to train on (e.g., law, code, math). Prior work has proposed a diverse set of methods to effici…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.