Showing cs.PFShow all
2 papers · 1 filter
cs.PF2025
Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
Youpeng Zhao, Jinpeng LV, Di Wu +2
Test-time scaling (TTS) has recently emerged as a promising direction to exploit the hidden reasoning capabilities of pre-trained large language models (LLMs). However, existing sc…
cs.PF2024
ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
Youpeng Zhao, Jun Wang
Large Language Models (LLMs) represent a revolutionary advancement in the contemporary landscape of artificial general intelligence (AGI). As exemplified by ChatGPT, LLM-based appl…