1 citations · 2 across the 15 of their papers we have counts for
23 papers
Causal Autoregressive Diffusion Language Model
Junhao Ruan, Bei Li, Yongjing Yin +6
In this work, we propose Causal Autoregressive Diffusion (CARD), a novel framework that unifies the training efficiency of ARMs with the high-throughput inference of diffusion mode…
Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text
Zhihao Xu, Rumei Li, Jiahuan Li +4
Enabling Large Language Models (LLMs) to effectively utilize tools in multi-turn interactions is essential for building capable autonomous agents. However, acquiring diverse and re…
Efficient Context Scaling with LongCat ZigZag Attention
Chen Zhang, Yang Bai, Jiahuan Li +19
We introduce LongCat ZigZag Attention (LoZA), which is a sparse attention scheme designed to transform any existing full-attention models into sparse versions with rather limited c…
A Preliminary Study on the Promises and Challenges of Native Top- Sparse Attention
Di Xiu, Hongyin Tang, Bolin Rong +4
Large Language Models (LLMs) are increasingly prevalent in the field of long-context modeling, however, their inference computational costs have become a critical bottleneck hinder…
A Survey on LLM Mid-Training
Chengying Tu, Xuemiao Zhang, Rongxiang Weng +6
Recent advances in foundation models have highlighted the significant benefits of multi-stage training, with a particular emphasis on the emergence of mid-training as a vital stage…
Autoformalizer with Tool Feedback
Qi Guo, Jianing Wang, Jianfei Zhang +8
Autoformalization addresses the scarcity of data for Automated Theorem Proving (ATP) by translating mathematical problems from natural language into formal statements. Efforts in r…