4 citations · 6 across the 7 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
Rui Pan, Yinwei Dai, Zhihao Zhang +3
Recent advances in inference-time compute have significantly improved performance on complex tasks by generating long chains of thought (CoTs) using Large Reasoning Models (LRMs).…
cs.LG2024★ 2 cited
METIS: Fast Quality-Aware RAG Systems with Configuration Adaptation
Siddhant Ray, Rui Pan, Zhuohan Gu +5
RAG (Retrieval Augmented Generation) allows LLMs (large language models) to generate better responses with external knowledge, but using more external knowledge often improves gene…
cs.LG2024
Optimizing Mixture-of-Experts Inference Time Combining Model Deployment and Communication Scheduling
Jialong Li, Shreyansh Tripathi, Lakshay Rastogi +3
As machine learning models scale in size and complexity, their computational requirements become a significant barrier. Mixture-of-Experts (MoE) models alleviate this issue by sele…