2 papers
cs.CL2026
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
Chaoyue He, Xin Zhou, Di Wang +3
Direct Preference Optimization (DPO) controls the trade-off between fitting preference labels and staying close to a reference model using a single global temperature beta, implici…
cs.MM2025
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
Lei Zhang, Xin Zhou, Chaoyue He +5
Environmental, Social, and Governance (ESG) reports are essential for evaluating sustainability practices, ensuring regulatory compliance, and promoting financial transparency. How…