7 papers
Multi-Adapter Representation Interventions via Energy Calibration
Manjiang Yu, Hongji Li, Junwei Chen +4
Representation intervention has emerged as a promising paradigm for aligning large language models toward desired behaviors without modifying model weights. Existing methods typica…
Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2
Zhanfeng Feng, Shuai Guo, Xin Di +3
This report describes Tail-Aware HiFloat4, our submission to the low-bit text-to-video generation quantization challenge. Our method adapts the public ViDiT-Q post-training quantiz…
Memento: Personalized RAG-Style Long-Retention Data Scaling for META Ads Recommendation
Xiaoyu Chen, Ruichen Wang, Jieming Di +21
Modeling of long history data suffers from long-context window attention dilution, system efficiency and catastrophic forgetting problems, where naive linear scaling approach like…
SLIM: Stealthy Low-Coverage Black-Box Watermarking via Latent-Space Confusion Zones
Hengyu Wu, Yang Cao
Training data is a critical and often proprietary asset in Large Language Model (LLM) development, motivating the use of data watermarking to embed model-transferable signals for u…
Null-Space Filtering for Data-Free Continual Model Merging: Preserving Stability, Promoting Plasticity
Zihuan Qiu, Lei Wang, Yang Cao +7
Data-free continual model merging (DFCMM) aims to fuse independently fine-tuned models into a single backbone that evolves with incoming tasks without accessing task data. This pap…
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
Jian Tian, Shuailong Li, Yang Cao +8
The evolution of Large Language Model (LLM) serving towards complex, distributed architectures--specifically the P/D-separated, large-scale DP+EP paradigm--introduces distinct sche…