activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding

Minjie Qiang, Mingming Zhang, Xiaoyi Bao +5

Foundation models have established unified representations for natural language processing, yet this paradigm remains largely unexplored for tabular data. Existing methods face fun…

cs.CL2026

Table as a Modality for Large Language Models

Liyao Li, Chao Ye, Wentao Ye +9

To migrate the remarkable successes of Large Language Models (LLMs), the community has made numerous efforts to generalize them to the table reasoning tasks for the widely deployed…

cs.CL2025

Chinese ModernBERT with Whole-Word Masking

Zeyu Zhao, Ningtao Wang, Xing Fu +1

Encoder-only Transformers have advanced along three axes -- architecture, data, and systems -- yielding Pareto gains in accuracy, speed, and memory efficiency. Yet these improvemen…

cs.CL2025

Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback

Kejin Liu, Junhong Lian, Xiang Ao +5

Accurate personalized headline generation hinges on precisely capturing user interests from historical behaviors. However, existing methods neglect personalized-irrelevant click no…

cs.CL2025

ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models

Hao Chen, Haoze Li, Zhiqing Xiao +6

Aligning general-purpose large language models (LLMs) to downstream tasks often incurs significant training adjustment costs. Prior research has explored various avenues to enhance…