5 papers
Binary Decompilation LLM with Feedback-Driven Multi-Turn Refinement
Peipei Liu, Jian Sun, Mingzhe Xing +5
Binary decompilation is fundamental to security tasks such as vulnerability discovery, malware inspection, and executable-only program understanding. Recent LLM-based decompilation…
Human-Less LLM Serving: Quantifying the Human Tax on Throughput
Jianhui Lian, Li Chen, Dan Li +1
Every major LLM serving system is designed to meet TTFT and TPOT SLOs. These metrics capture latency as a human user perceives it, and the mechanisms built to satisfy them are now…
Supercharging Packet-level Network Simulation of Large Model Training via Memoization and Fast-Forwarding
Fei Long, Kaihui Gao, Li Chen +6
Packet-level discrete-event simulation (PLDES) is a prevalent tool for evaluating detailed performance of large model training. Although PLDES offers high fidelity and generality,…
A Large-Scale IPv6-Based Measurement of the Starlink Network
Bingsen Wang, Xiaohui Zhang, Shuai Wang +4
Low Earth Orbit (LEO) satellite networks have attracted considerable attention for their ability to deliver global, low-latency broadband Internet services. In this paper, we prese…
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
Zhiyuan Wu, Shuai Wang, Li Chen +5
Video diffusion models (VDMs) perform attention computation over the 3D spatio-temporal domain. Compared to large language models (LLMs) processing 1D sequences, their memory consu…