3 papers
cs.LG2025
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
Zhengze Zhang, Shiqi Wang, Yiqun Shen +5
Large language models (LLMs) have demonstrated exceptional performance across various applications, but their conversational abilities decline sharply as model size decreases, pres…
cs.LG2025
Corporate Fraud Detection in Rich-yet-Noisy Financial Graph
Shiqi Wang, Zhibo Zhang, Libing Fang +2
Corporate fraud detection aims to automatically recognize companies that conduct wrongful activities such as fraudulent financial statements or illegal insider trading. Previous le…
cs.CL2024
Reward Difference Optimization For Sample Reweighting In Offline RLHF
Shiqi Wang, Zhengze Zhang, Rui Zhao +2
With the rapid advances in Large Language Models (LLMs), aligning LLMs with human preferences become increasingly important. Although Reinforcement Learning with Human Feedback (RL…