1 citations · 1 across the 8 of their papers we have counts for
6 papers · 1 filter
CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution
Zihao Ye, Yingyi Huang, Hongyi Jin +11
GPU kernel agents and GPU programming languages have advanced separately, leaving expert kernels difficult to reproduce. Agents usually treat the compiler as a fixed black box and…
Reasoning Quality Emerges Early: Data Curation for Reasoning Models
Hongyi Henry Jin, Wenhan Yang, Meysam Ghaffari +2
Supervised fine-tuning (SFT) on a small, high-quality set of long reasoning traces is an effective approach for eliciting strong reasoning capabilities in Large Language Models (LL…
Discretization-free Multicalibration through Loss Minimization over Tree Ensembles
Hongyi Henry Jin, Zijun Ding, Dung Daniel Ngo +1
In recent years, multicalibration has emerged as a desirable learning objective for ensuring that a predictor is calibrated across a rich collection of overlapping subpopulations.…
WebLLM: A High-Performance In-Browser LLM Inference Engine
Charlie F. Ruan, Yucheng Qin, Akaash R. Parthasarathy +11
Advancements in large language models (LLMs) have unlocked remarkable capabilities. While deploying these models typically requires server-grade GPUs and cloud-based inference, the…
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang +4
In the rapidly evolving landscape of artificial intelligence (AI), generative large language models (LLMs) stand at the forefront, revolutionizing how we interact with our data. Ho…
Relax: Composable Abstractions for End-to-End Dynamic Machine Learning
Ruihang Lai, Junru Shao, Siyuan Feng +16
Dynamic shape computations have become critical in modern machine learning workloads, especially in emerging large language models. The success of these models has driven the deman…