5 papers
MASTE: A Multi-Agent Pipeline for Zero-Shot Aspect Sentiment Triplet Extraction
Ao Hong, Lehang Wang, Zhirun Yue +3
Aspect Sentiment Triplet Extraction (ASTE) requires jointly identifying (aspect, opinion, sentiment) triples from a given review sentence. While large language models (LLMs) achiev…
Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning
Kai Qin, Jiaqi Wu, Jianxiang He +8
As Large Language Models (LLMs) demonstrate remarkable capabilities learned from vast corpora, concerns regarding data privacy and safety are receiving increasing attention. LLM un…
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework
Kai Qin, Liangxin Liu, Yu Liang +7
Reward Models (RMs) are critical components in the Reinforcement Learning from Human Feedback (RLHF) pipeline, directly determining the alignment quality of Large Language Models (…
Let the Agent Search: Autonomous Exploration Beats Rigid Workflows in Temporal Question Answering
Xufei Lv, Jiahui Yang, Haoyuan Sun +5
Temporal Knowledge Graph Question Answering (TKGQA) is challenging because it requires multi-hop reasoning under complex temporal constraints. Recent LLM-based approaches have impr…
The Hidden Link Between RLHF and Contrastive Learning
Xufei Lv, Kehai Chen, Haoyuan Sun +3
Alignment of large language models (LLMs) with human values has recently garnered significant attention, with prominent examples including the canonical yet costly Reinforcement Le…