5 papers · 1 filter
Enhancing Judgment Document Generation via Agentic Legal Information Collection and Rubric-Guided Optimization
Weihang Su, Xuanyi Chen, Yueyue Wu +2
Automating the drafting of judgment documents is pivotal to judicial efficiency, yet it remains challenging due to the dual requirements of comprehensive retrieval of legal informa…
On the Role of Preference Variance in Preference Optimization
Jiacheng Guo, Zihao Li, Jiahao Qiu +2
Direct Preference Optimization (DPO) has emerged as an important approach for learning from human preferences in aligning large language models (LLMs). However, collecting human pr…
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
Jiahao Qiu, Yifu Lu, Yifan Zeng +9
Inference-time alignment enhances the performance of large language models without requiring additional training or fine-tuning but presents challenges due to balancing computation…
Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data?
Xuan Qi, Jiahao Qiu, Xinzhe Juan +2
Aligning large language models (LLMs) with human preferences remains a key challenge in AI. Preference-based optimization methods, such as Reinforcement Learning with Human Feedbac…
Temporal Consistency for LLM Reasoning Process Error Identification
Jiacheng Guo, Yue Wu, Jiahao Qiu +4
Verification is crucial for effective mathematical reasoning. We present a new temporal consistency method where verifiers iteratively refine their judgments based on the previous…