From the 1 of 12 linked papers with an AI index.
12 papers
Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways
Shuyi Miao, Wangjie Qiu, Pengyang Shao +4
Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mech…
CoLAS: Multimodal Corroboration of Latent Asset Signals for Financial Trading
Yanzheng Jin, Pengyang Shao, Xiaohao Liu +3
CoLAS is a multimodal learning framework that extracts trading signals by identifying and reinforcing shared information across price data, news, and sentiment, improving robustnes…
BalDRO: A Distributionally Robust Optimization based Framework for Large Language Model Unlearning
Pengyang Shao, Naixin Zhai, Lei Chen +4
As Large Language Models (LLMs) increasingly shape online content, removing targeted information from well-trained LLMs (also known as LLM unlearning) has become critical for web g…
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
Jilong Liu, Yonghui Yang, Pengyang Shao +5
Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes training sensitive to imperfect sup…
Sharpness-Aware Poisoning: Enhancing Transferability of Injective Attacks on Recommender Systems
Junsong Xie, Yonghui Yang, Pengyang Shao +1
Recommender Systems~(RS) have been shown to be vulnerable to injective attacks, where attackers inject limited fake user profiles to promote the exposure of target items to real us…
Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning
Naixin Zhai, Pengyang Shao, Binbin Zheng +4
Machine unlearning aims to forget sensitive knowledge from Large Language Models (LLMs) while maintaining general utility. However, existing approaches typically treat all tokens i…