6 papers
Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
Wenzhe Xu, Biao Liu, Yiyang Sun +2
Multi-Objective Alignment aims to align Large Language Models (LLMs) with diverse and often conflicting human values by optimizing multiple objectives simultaneously. Existing meth…
VRM: Teaching Reward Models to Understand Authentic Human Preferences
Biao Liu, Ning Xu, Junming Yang +2
Large Language Models (LLMs) have achieved remarkable success across diverse natural language tasks, yet the reward models employed for aligning LLMs often encounter challenges of…
iScript: A Domain-Adapted Large Language Model and Benchmark for Physical Design Tcl Script Generation
Ning Xu, Zhaoyang Zhang, Senlin Shu +10
Modern EDA flows rely heavily on Tcl scripting, yet general LLMs perform poorly in this domain due to extreme data scarcity, domain-specific semantics, and the high reliability req…
Enriching Knowledge Distillation with Intra-Class Contrastive Learning
Hua Yuan, Ning Xu, Xin Geng +1
Since the advent of knowledge distillation, much research has focused on how the soft labels generated by the teacher model can be utilized effectively. Existing studies points out…
Towards Understanding Feature Learning in Parameter Transfer
Hua Yuan, Xuran Meng, Qiufeng Wang +6
Parameter transfer is a central paradigm in transfer learning, enabling knowledge reuse across tasks and domains by sharing model parameters between upstream and downstream models.…
Reduction-based Pseudo-label Generation for Instance-dependent Partial Label Learning
Congyu Qiao, Ning Xu, Yihao Hu +1
Instance-dependent Partial Label Learning (ID-PLL) aims to learn a multi-class predictive model given training instances annotated with candidate labels related to features, among…