10 papers
PrunePath: Towards Highly Structured Sparse Language Models
Zhexuan Gu, Zixun Fu, Yancheng Yuan
Feed-forward networks (FFNs) dominate the parameter count and computation of modern language models, yet existing pruning methods often struggle to convert sparsity into hardware-f…
Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning
Chenlu Ding, Jiancan Wu, Yanchen Luo +3
Large language models (LLMs) often fail to reason under temporal cutoffs: when prompted to answer from the standpoint of an earlier time, they exploit knowledge that became availab…
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
Zonglin Yang, Zhexuan Gu, Yancheng Yuan
Training a deep neural network with the outputs of selected layers satisfying linear constraints is required in many contemporary data-driven applications. While this can be achiev…
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
Chenlu Ding, Jiancan Wu, Leheng Sheng +4
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities across vision-language tasks, yet their large-scale deployment raises pressing concerns about mem…
Delayed Feedback Modeling with Influence Functions
Chenlu Ding, Jiancan Wu, Yancheng Yuan +5
In online advertising under the cost-per-conversion (CPA) model, accurate conversion rate (CVR) prediction is crucial. A major challenge is delayed feedback, where conversions may…
Addressing Missing Data Issue for Diffusion-based Recommendation
Wenyu Mao, Zhengyi Yang, Jiancan Wu +4
Diffusion models have shown significant potential in generating oracle items that best match user preference with guidance from user historical interaction sequences. However, the…