5 papers
Process Rewards with Learned Reliability
Jinyuan Li, Langlin Huang, Chengsong Huang +5
Process Reward Models (PRMs) provide step-level feedback for reasoning, but current PRMs usually output only a single reward score for each step. Downstream methods must therefore…
TabDLM: Free-Form Tabular Data Generation via Joint Numerical-Language Diffusion
Donghong Cai, Jiarui Feng, Yanbo Wang +3
Synthetic tabular data generation has attracted growing attention due to its importance for data augmentation, foundation models, and privacy. However, real-world tabular datasets…
Nonsense Helps: Prompt Space Perturbation Broadens Reasoning Exploration
Langlin Huang, Chengsong Huang, Jinyuan Li +3
Reinforcement learning with verifiable rewards, particularly Group Relative Policy Optimization (GRPO), has significantly advanced the reasoning capabilities of Large Language Mode…
GRIP: In-Parameter Graph Reasoning through Fine-Tuning Large Language Models
Jiarui Feng, Donghong Cai, Yixin Chen +1
Large Language Models (LLMs) have demonstrated remarkable capabilities in modeling sequential textual data and generalizing across diverse tasks. However, effectively adapting LLMs…
Addressing accuracy and hallucination of LLMs in Alzheimer's disease research through knowledge graphs
Tingxuan Xu, Jiarui Feng, Justin Melendez +6
In the past two years, large language model (LLM)-based chatbots, such as ChatGPT, have revolutionized various domains by enabling diverse task completion and question-answering ca…