9 papers
Generalization in Federated Learning: A Conditional Mutual Information Framework
Ziqiao Wang, Cheng Long, Yongyi Mao
Federated learning (FL) is a widely adopted privacy-preserving distributed learning framework, yet its generalization performance remains less explored compared to centralized lear…
Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
Wei Yao, Wenkai Yang, Ziqiao Wang +2
As large language models advance toward superhuman performance, ensuring their alignment with human values and abilities grows increasingly complex. Weak-to-strong generalization o…
The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration
Wei Yao, Wenkai Yang, Gengze Xu +3
Weak-to-strong generalization, where weakly supervised strong models outperform their weaker teachers, offers a promising approach to aligning superhuman models with human values.…
Distributional Information Embedding: A Framework for Multi-bit Watermarking
Haiyun He, Yepeng Liu, Ziqiao Wang +2
This paper introduces a novel problem, distributional information embedding, motivated by the practical demands of multi-bit watermarking for large language models (LLMs). Unlike t…
LH-Mix: Local Hierarchy Correlation Guided Mixup over Hierarchical Prompt Tuning
Fanshuang Kong, Richong Zhang, Ziqiao Wang
Hierarchical text classification (HTC) aims to assign one or more labels in the hierarchy for each text. Many methods represent this structure as a global hierarchy, leading to red…
MOMA: Masked Orthogonal Matrix Alignment for Zero-Additional-Parameter Model Merging
Fanshuang Kong, Richong Zhang, Zhijie Nie +4
Model merging offers a scalable alternative to multi-task learning but often yields suboptimal performance on classification tasks. We attribute this degradation to a geometric mis…