1 citations · 1 across the 17 of their papers we have counts for
7 papers · 1 filter
Zero-Shot Transfer Capabilities of the Sundial Foundation Model for Leaf Area Index Forecasting
Peining Zhang, Hongchen Qin, Haochen Zhang +3
This work investigates the zero-shot forecasting capability of time series foundation models for Leaf Area Index (LAI) forecasting in agricultural monitoring. Using the HiQ dataset…
Support Basis: Fast Attention Beyond Bounded Entries
Maryam Aliakbarpour, Vladimir Braverman, Junze Yin +1
Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks. However, the quadratic complexity of softmax attention remains a central bottlen…
Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding
StepFun, :, Bin Wang +195
Large language models (LLMs) face low hardware efficiency during decoding, especially for long-context reasoning tasks. This paper introduces Step-3, a 321B-parameter VLM with hard…
Few-Shot Learning for Industrial Time Series: A Comparative Analysis Using the Example of Screw-Fastening Process Monitoring
Xinyuan Tu
Few-shot learning (FSL) has shown promise in vision but remains largely unexplored for \emph{industrial} time-series data, where annotating every new defect is prohibitively expens…
Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM Pretraining
Haochen Zhang, Junze Yin, Guanchu Wang +5
Low-rank optimization has emerged as a promising approach to enabling memory-efficient training of large language models (LLMs). Existing low-rank optimization methods typically pr…
Statistical Guarantees for Lifelong Reinforcement Learning using PAC-Bayes Theory
Zhi Zhang, Chris Chow, Yasi Zhang +7
Lifelong reinforcement learning (RL) has been developed as a paradigm for extending single-task RL to more realistic, dynamic settings. In lifelong RL, the "life" of an RL agent is…