3 citations · 6 across the 7 of their papers we have counts for
6 papers · 1 filter
Theoretical Foundation of Flow-Based Time Series Generation: Provable Approximation, Generalization, and Efficiency
Jiangxuan Long, Zhao Song, Chiwun Yang
Recent studies suggest utilizing generative models instead of traditional auto-regressive algorithms for time series forecasting (TSF) tasks. These non-auto-regressive approaches i…
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
Majid Daliri, Zhao Song, Chiwun Yang
Recently, 1-bit Large Language Models (LLMs) have emerged, showcasing an impressive combination of efficiency and performance that rivals traditional LLMs. Research by Wang et al.…
A Theoretical Insight into Attack and Defense of Gradient Leakage in Transformer
Chenyang Li, Zhao Song, Weixin Wang +1
The Deep Leakage from Gradient (DLG) attack has emerged as a prevalent and highly effective method for extracting sensitive training data by inspecting exchanged gradients. This ap…
One Pass Streaming Algorithm for Super Long Token Attention Approximation in Sublinear Space
Raghav Addanki, Chenyang Li, Zhao Song +1
Attention computation takes both the time complexity of and the space complexity of simultaneously, which makes deploying Large Language Models (LLMs) in streamin…
An Automatic Learning Rate Schedule Algorithm for Achieving Faster Convergence and Steeper Descent
Zhao Song, Chiwun Yang
The delta-bar-delta algorithm is recognized as a learning rate adaptation technique that enhances the convergence speed of the training process in optimization by dynamically sched…
Fine-tune Language Models to Approximate Unbiased In-context Learning
Timothy Chu, Zhao Song, Chiwun Yang
In-context learning (ICL) is an astonishing emergent ability of large language models (LLMs). By presenting a prompt that includes multiple input-output pairs as examples and intro…