4 papers
Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
Bolian Li, Yanran Wu, Xinyu Luo +1
Aligning large language models (LLMs) with human preferences has become a critical step in their development. Recent research has increasingly focused on test-time alignment, where…
Not All Water Consumption Is Equal: A Water Stress Weighted Metric for Sustainable Computing
Yanran Wu, Inez Hua, Yi Ding
Water consumption is an increasingly critical dimension of computing sustainability, especially as AI workloads rapidly scale. However, current water impact assessment often overlo…
Unveiling Environmental Impacts of Large Language Model Serving: A Functional Unit View
Yanran Wu, Inez Hua, Yi Ding
Large language models (LLMs) offer powerful capabilities but come with significant environmental impact, particularly in carbon emissions. Existing studies benchmark carbon emissio…
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
Tianyao Shi, Yanran Wu, Sihang Liu +1
LLMs have been widely adopted across many real-world applications. However, their widespread use comes with significant environmental costs due to their high computational intensit…