7 papers
Borrowing from anything: A generalizable framework for reference-guided instance editing
Shengxiao Zhou, Chenghua Li, Jianhao Huang +2
Reference-guided instance editing is fundamentally limited by semantic entanglement, where a reference's intrinsic appearance is intertwined with its extrinsic attributes. The key…
Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from -Parity
Jianhao Huang, Baharan Mirzasoleiman
Masked Diffusion Language Models have recently emerged as a powerful generative paradigm, yet their generalization properties remain understudied compared to their auto-regressive…
Stability-Weighted Decoding for Diffusion Language Models
Yue Wu, Jian Huang
Diffusion large language models (dLLMs) enable parallel text generation by iteratively denoising a fully masked sequence, unmasking a subset of masked tokens at each step. Existing…
How Transformers Learn to Plan via Multi-Token Prediction
Jianhao Huang, Zhanpeng Zhou, Renqiu Xia +3
While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global structure in reasoning tasks. Multi-token predi…
Beyond What Seems Necessary: Hidden Gains from Scaling Training-Time Reasoning Length under Outcome Supervision
Yihao Xue, Allan Zhang, Jianhao Huang +2
Training LLMs to think and reason for longer has become a key ingredient in building state-of-the-art models that can solve complex problems previously out of reach. Recent efforts…
SL-ACC: A Communication-Efficient Split Learning Framework with Adaptive Channel-wise Compression
Zehang Lin, Zheng Lin, Miao Yang +7
The increasing complexity of neural networks poses a significant barrier to the deployment of distributed machine learning (ML) on resource-constrained devices, such as federated l…