3 papers
cs.LG2026
Intrinsic Mutual Information as a Modulator for Preference Optimization
Peng Liao, Peijia Zheng, Lingbo Li +2
Offline preference optimization methods, such as Direct Preference Optimization (DPO), offer significant advantages in aligning Large Language Models (LLMs) with human values. Howe…
cs.CL2026
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
Weibo Zhou, Lingbo Li, Shangsong Liang
The scarcity and high cost of high-quality domain-specific question-answering (QA) datasets limit supervised fine-tuning of large language models (LLMs). We introduce $\textbf{D-SC…
cs.CL2025
LETToT: Label-Free Evaluation of Large Language Models On Tourism Using Expert Tree-of-Thought
Ruiyan Qi, Congding Wen, Weibo Zhou +3
Evaluating large language models (LLMs) in specific domain like tourism remains challenging due to the prohibitive cost of annotated benchmarks and persistent issues like hallucina…