collaborators

7 papers

cs.LG2026

DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs

Zeyu Cao, Xuan Guo, Cheng Zhang +3

As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets. This paper investigates whether these retired GPUs can find a produ…

cs.AR2026

KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware

Jiayi Nie, Haoran Wu, Yao Lai +9

New AI accelerators with novel instruction set architectures (ISAs) often require developers to manually craft low-level kernels, a time-consuming and error-prone process that does…

cs.AI2026

Reasoning Compression with Mixed-Policy Distillation

Han Yang, Mingyan Wu, Bailan He +4

Reasoning-centric large language models (LLMs) achieve strong performance by generating intermediate reasoning trajectories, but often incur excessive token usage and high inferenc…

cs.AI2025

Clinical-R1: Empowering Large Language Models for Faithful and Comprehensive Reasoning with Clinical Objective Relative Policy Optimization

Boyang Gu, Hongjian Zhou, Bradley Max Segal +6

Recent advances in large language models (LLMs) have shown strong reasoning capabilities through large-scale pretraining and post-training reinforcement learning, demonstrated by D…

cs.AI2025

Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning

Zifeng Ding, Shenyang Huang, Zeyu Cao +11

Forecasting future links is a central task in temporal graph (TG) reasoning, requiring models to leverage historical interactions to predict upcoming ones. Traditional neural appro…

cs.CL2025

Scaling Laws For Mixed Quantization

Zeyu Cao, Boyang Gu, Cheng Zhang +5

Post-training quantization of Large Language Models (LLMs) has proven effective in reducing the memory and computational requirements for inference. In this study, we focus on a st…