2 papers
cs.CL2026
Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation
Mingqi Gao, Anthony Sicilia, Weiyan Shi
Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inferenc…
cs.LG2026
Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding
Weiye Shi, Fanxu Meng, Muhan Zhang
Multi-head latent attention (MLA) is increasingly important for long-context LLM inference because compact latent states replace the growing key-value (KV) cache and reduce decodin…