3 papers
cs.CL2026
TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning
Muyu Pan, Shu Zhao, Nan Zhang +4
This paper investigates large language model (LLM) abstention learning, specifically using ternary reward, which incentivize truthfulness in large language models. This paper exten…
cs.CL2025
LangSAT: A Novel Framework Combining NLP and Reinforcement Learning for SAT Solving
Muyu Pan, Matthew Walter, Dheeraj Kodakandla +1
Our work presents a novel reinforcement learning (RL) based framework to optimize heuristic selection within the conflict-driven clause learning (CDCL) process, improving the effic…
cs.CL2025
Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic
Muyu Pan, Dheeraj Kodakandla, Mahfuza Farooque
Recent advances in natural language processing (NLP), particularly large language models (LLMs), have motivated the automatic translation of natural language statements into formal…