collaborators

5 papers

cs.CL20261 cited

QFT: Quantized Full-parameter Tuning of LLMs with Affordable Resources

Zhikai Li, Xiaoxuan Liu, Banghua Zhu +3

Large Language Models (LLMs) have showcased remarkable impacts across a wide spectrum of natural language processing tasks. Fine-tuning these pretrained models on downstream datase…

cs.AI2026

: Faster Test-Time Scaling through Speculative Drafts

Mert Cemri, Nived Rajaraman, Rishabh Tiwari +6

Scaling test-time compute has driven the recent advances in the reasoning capabilities of large language models (LLMs), typically by allocating additional computation for more thor…

cs.AI2025

GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models

Yue Zhang, Jiaxin Zhang, Qiuyu Ren +5

We introduce \textbf{GAUSS} (\textbf{G}eneral \textbf{A}ssessment of \textbf{U}nderlying \textbf{S}tructured \textbf{S}kills in Mathematics), a benchmark that evaluates LLMs' mathe…

cs.AI2025

WGRAMMAR: Leverage Prior Knowledge to Accelerate Structured Decoding

Ran Wang, Xiaoxuan Liu, Hao Ren +3

Structured decoding enables large language models (LLMs) to generate outputs in formats required by downstream systems, such as HTML or JSON. However, existing methods suffer from…

cs.CL2025

Sequential Diagnosis with Language Models

Harsha Nori, Mayank Daswani, Christopher Kelly +12

Artificial intelligence holds great promise for expanding access to expert medical knowledge and reasoning. However, most evaluations of language models rely on static vignettes an…