most citedXGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models

8 citations · 8 across the 6 of their papers we have counts for

collaborators

6 papers

cs.AI2026

FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems

Shanli Xing, Yiyan Zhai, Alexander Jiang +10

Recent advances show that large language models (LLMs) can act as autonomous agents capable of generating GPU kernels, but integrating these AI-generated kernels into real-world in…

cs.DC2024

A System for Microserving of LLMs

Hongyi Jin, Ruihang Lai, Charlie F. Ruan +5

The recent advances in LLMs bring a strong demand for efficient system support to improve overall serving efficiency. As LLM inference scales towards multiple GPUs and even multipl…

cs.LG2024

WebLLM: A High-Performance In-Browser LLM Inference Engine

Charlie F. Ruan, Yucheng Qin, Akaash R. Parthasarathy +11

Advancements in large language models (LLMs) have unlocked remarkable capabilities. While deploying these models typically requires server-grade GPUs and cloud-based inference, the…

cs.SD2024

Local deployment of large-scale music AI models on commodity hardware

Xun Zhou, Charlie Ruan, Zihe Zhao +2

We present the MIDInfinite, a web application capable of generating symbolic music using a large-scale generative AI model locally on commodity hardware. Creating this demo involve…

cs.CL2024★ 8 cited

XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models

Yixin Dong, Charlie F. Ruan, Yaxing Cai +4

The applications of LLM Agents are becoming increasingly complex and diverse, leading to a high demand for structured outputs that can be parsed into code, structured function call…

cs.SE2024

Productively Deploying Emerging Models on Emerging Platforms: A Top-Down Approach for Testing and Debugging

Siyuan Feng, Jiawei Liu, Ruihang Lai +4

While existing machine learning (ML) frameworks focus on established platforms, like running CUDA on server-grade GPUs, there have been growing demands to enable emerging AI applic…