activity
20242026
collaborators

6 papers

cs.CL2026

TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding

Minjie Qiang, Mingming Zhang, Xiaoyi Bao +5

Foundation models have established unified representations for natural language processing, yet this paradigm remains largely unexplored for tabular data. Existing methods face fun…

cs.LG2026

KMLP: A Scalable Hybrid Architecture for Web-Scale Tabular Data Modeling

Mingming Zhang, Pengfei Shi, Zhiqing Xiao +8

Predictive modeling on web-scale tabular data with billions of instances and hundreds of heterogeneous numerical features faces significant scalability challenges. These features e…

cs.CL2025

Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback

Kejin Liu, Junhong Lian, Xiang Ao +5

Accurate personalized headline generation hinges on precisely capturing user interests from historical behaviors. However, existing methods neglect personalized-irrelevant click no…

cs.LG2025

Beyond Tree Models: A Hybrid Model of KAN and gMLP for Large-Scale Financial Tabular Data

Mingming Zhang, Jiahao Hu, Pengfei Shi +8

Tabular data plays a critical role in real-world financial scenarios. Traditionally, tree models have dominated in handling tabular data. However, financial datasets in the industr…

cs.AI2024

AIGT: AI Generative Table Based on Prompt

Mingming Zhang, Zhiqing Xiao, Guoshan Lu +5

Tabular data, which accounts for over 80% of enterprise data assets, is vital in various fields. With growing concerns about privacy protection and data-sharing restrictions, gener…

cs.LG2024

Estimating Conditional Average Treatment Effects via Sufficient Representation Learning

Pengfei Shi, Wei Zhong, Xinyu Zhang +4

Estimating the conditional average treatment effects (CATE) is very important in causal inference and has a wide range of applications across many fields. In the estimation process…