3 papers
cs.CL2026
Fixing the Broken Compass: Diagnosing and Improving Inference-Time Reward Modeling
Jiachun Li, Pengfei Cao, Zhuoran Jin +6
Inference-time scaling techniques have shown promise in enhancing the reasoning capabilities of large language models (LLMs). While recent research has primarily focused on trainin…
cs.CL2024
AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation
Zhitao He, Pengfei Cao, Chenhao Wang +7
With the development of deep learning, natural language processing technology has effectively improved the efficiency of various aspects of the traditional judicial industry. Howev…
cs.CL2024
Cutting Off the Head Ends the Conflict: A Mechanism for Interpreting and Mitigating Knowledge Conflicts in Language Models
Zhuoran Jin, Pengfei Cao, Hongbang Yuan +6
Recently, retrieval augmentation and tool augmentation have demonstrated a remarkable capability to expand the internal memory boundaries of language models (LMs) by providing exte…