9 papers
TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload
Zhiben Chen, Youpeng Zhao, Yang Sui +2
Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive (AR) models, offering better hardware utilization and bidirectional context thro…
A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
Yuanhao Zou, Shengji Jin, Andong Deng +3
Effectively applying Vision-Language Models (VLMs) to Video Question Answering (VideoQA) hinges on selecting a concise yet comprehensive set of frames, as processing entire videos…
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
Shakya Jayakody, Youpeng Zhao, Chinmay Dhanraj Nehate +1
The rise of million-token, agent-based applications has placed unprecedented demands on large language model (LLM) inference services. The long-running nature of these tasks increa…
LMSeg: Unleashing the Power of Large-Scale Models for Open-Vocabulary Semantic Segmentation
Huadong Tang, Youpeng Zhao, Yan Huang +3
It is widely agreed that open-vocabulary-based approaches outperform classical closed-set training solutions for recognizing unseen objects in images for semantic segmentation. Exi…
Classifier Enhancement Using Extended Context and Domain Experts for Semantic Segmentation
Huadong Tang, Youpeng Zhao, Min Xu +2
Prevalent semantic segmentation methods generally adopt a vanilla classifier to categorize each pixel into specific classes. Although such a classifier learns global information fr…
Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
Youpeng Zhao, Jinpeng LV, Di Wu +2
Test-time scaling (TTS) has recently emerged as a promising direction to exploit the hidden reasoning capabilities of pre-trained large language models (LLMs). However, existing sc…