activity
20242026
collaborators

8 papers

cs.DC2026

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training

Tan Zhiqiang, Zhiqiang Tan, Maoxin Wang +11

It is well established that the reasoning capabilities of large language models (LLMs) can be improved by applying reinforcement learning (RL) in a post-training stage. In a standa…

cs.CE2026

Position: LLM Inference Should Be Evaluated as Energy-to-Token Production

Xiang Liu, Shimiao Yuan, Zhenheng Tang +5

LLM inference is still evaluated mainly as a model or software problem: accuracy, latency, throughput, and hardware utilization. This is incomplete. At deployment scale, the releva…

cs.DC2025

Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis

Weile Luo, Ruibo Fan, Zeyu Li +4

This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper'…

cs.CV2025

RA-NeRF: Robust Neural Radiance Field Reconstruction with Accurate Camera Pose Estimation under Complex Trajectories

Qingsong Yan, Qiang Wang, Kaiyong Zhao +4

Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have emerged as powerful tools for 3D reconstruction and SLAM tasks. However, their performance depends heavily on ac…

cs.DC2025

BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Yuxin Wang, Yuhan Chen, Zeyu Li +11

Serving systems for Large Language Models (LLMs) are often optimized to improve quality of service (QoS) and throughput. However, due to the lack of open-source LLM serving workloa…

cs.CV2025

SphereFusion: Efficient Panorama Depth Estimation via Gated Fusion

Qingsong Yan, Qiang Wang, Kaiyong Zhao +4

Due to the rapid development of panorama cameras, the task of estimating panorama depth has attracted significant attention from the computer vision community, especially in applic…