1 citations · 1 across the 4 of their papers we have counts for
4 papers
EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
Zhiye Song, Kyungmi Lee, Eun Kyung Lee +3
We present EnergyLens, an end-to-end framework for energy-aware large language model (LLM) inference optimization. As LLMs scale, predicting and reducing their energy footprint has…
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
Kyungmi Lee, Zhiye Song, Eun Kyung Lee +3
As AI workloads drive increases in datacenter power consumption, accurate GPU power estimation is critical for proactive power management. However, existing power models face a sca…
GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers
Zhiye Song, Steve Dai, Ben Keller +1
Diffusion models have revolutionized video generation, becoming essential tools in creative content generation and physical simulation. Transformer-based architectures (DiTs) and c…
Striped Attention: Faster Ring Attention for Causal Transformers
William Brandon, Aniruddha Nrusimha, Kevin Qian +4
To help address the growing demand for ever-longer sequence lengths in transformer models, Liu et al. recently proposed Ring Attention, an exact attention algorithm capable of over…