9 papers
Fast KV Compaction via Attention Matching
Adam Zweiger, Xinghong Fu, Han Guo +1
Scaling language models to long contexts is often bottlenecked by the size of the key-value (KV) cache. In deployed settings, long contexts are typically managed through compaction…
Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting
Xinghong Fu, Yanhong Li, Georgios Papaioannou +1
Learning time series foundation models has been shown to be a promising approach for zero-shot time series forecasting across diverse time series domains. Insofar as scaling has be…
Query as Anchor: Scenario-Adaptive User Representation via Large Language Model
Jiahao Yuan, Yike Xu, Jinyong Wen +9
Industrial-scale user representation learning requires balancing robust universality with acute task-sensitivity. However, existing paradigms primarily yield static, task-agnostic…
Table as a Modality for Large Language Models
Liyao Li, Chao Ye, Wentao Ye +9
To migrate the remarkable successes of Large Language Models (LLMs), the community has made numerous efforts to generalize them to the table reasoning tasks for the widely deployed…
Chinese ModernBERT with Whole-Word Masking
Zeyu Zhao, Ningtao Wang, Xing Fu +1
Encoder-only Transformers have advanced along three axes -- architecture, data, and systems -- yielding Pareto gains in accuracy, speed, and memory efficiency. Yet these improvemen…
ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models
Hao Chen, Haoze Li, Zhiqing Xiao +6
Aligning general-purpose large language models (LLMs) to downstream tasks often incurs significant training adjustment costs. Prior research has explored various avenues to enhance…