5 papers · 1 filter
LLaDA2.1: Speeding Up Text Diffusion via Token Editing
Tiwei Bie, Maosong Cao, Xiang Cao +47
While LLaDA2.0 showcased the scaling potential of 100B-level block-diffusion models and their inherent parallelization, the delicate equilibrium between decoding speed and generati…
Per-parameter Task Arithmetic for Unlearning in Large Language Models
Chengyi Cai, Zesheng Ye, Jiangchao Yao +5
In large language model (LLM) unlearning, private information is required to be removed. Task arithmetic unlearns by subtracting a specific task vector (TV)--defined as the paramet…
TableGPT-R1: Advancing Tabular Reasoning Through Reinforcement Learning
Saisai Yang, Qingyi Huang, Jing Yuan +13
Tabular data serves as the backbone of modern data analysis and scientific research. While Large Language Models (LLMs) fine-tuned via Supervised Fine-Tuning (SFT) have significant…
LLaDA2.0: Scaling Up Diffusion Language Models to 100B
Tiwei Bie, Maosong Cao, Kun Chen +28
This paper presents LLaDA2.0 -- a tuple of discrete diffusion large language models (dLLM) scaling up to 100B total parameters through systematic conversion from auto-regressive (A…
TableGPT2: A Large Multimodal Model with Tabular Data Integration
Aofeng Su, Aowen Wang, Chao Ye +30
The emergence of models like GPTs, Claude, LLaMA, and Qwen has reshaped AI applications, presenting vast new opportunities across industries. Yet, the integration of tabular data r…