9 papers
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
Hao Chen, Qi Zhang, Liyao Li +7
Adapting Large Language Models (LLMs) to specialized domains typically incurs high data and computational overhead. While prior efficiency efforts have largely treated data selecti…
TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding
Minjie Qiang, Mingming Zhang, Xiaoyi Bao +5
Foundation models have established unified representations for natural language processing, yet this paradigm remains largely unexplored for tabular data. Existing methods face fun…
KMLP: A Scalable Hybrid Architecture for Web-Scale Tabular Data Modeling
Mingming Zhang, Pengfei Shi, Zhiqing Xiao +8
Predictive modeling on web-scale tabular data with billions of instances and hundreds of heterogeneous numerical features faces significant scalability challenges. These features e…
Table as a Modality for Large Language Models
Liyao Li, Chao Ye, Wentao Ye +9
To migrate the remarkable successes of Large Language Models (LLMs), the community has made numerous efforts to generalize them to the table reasoning tasks for the widely deployed…
Chinese ModernBERT with Whole-Word Masking
Zeyu Zhao, Ningtao Wang, Xing Fu +1
Encoder-only Transformers have advanced along three axes -- architecture, data, and systems -- yielding Pareto gains in accuracy, speed, and memory efficiency. Yet these improvemen…
Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback
Kejin Liu, Junhong Lian, Xiang Ao +5
Accurate personalized headline generation hinges on precisely capturing user interests from historical behaviors. However, existing methods neglect personalized-irrelevant click no…