4 papers
Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference
Yuansheng Chen, Yue Zhang, Xuan Mo +2
With the rapid growth of interactive applications in large language model (LLM) online services, maintaining high system throughput while ensuring user-perceived latency has become…
UNIVID: Unified Vision-Language Model for Video Moderation
Kejuan Yang, Yizhuo Zhang, Mingyuan Du +6
Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to support downstream enforcement. Tr…
DeployBench: Benchmarking LLM Agents for Research Artifact Deployment
Yuanli Wang, Yaoyao Qian, Yue Zhang +8
LLM agents have made rapid progress on software engineering and ML research tasks, but these advances often assume access to a working runnable environment. For research artifacts…
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
Rui Jiao, Yue Zhang, Jinku Li
We present a novel framework addressing a critical vulnerability in Large Language Models (LLMs): the prevalence of factual inaccuracies within intermediate reasoning steps despite…