1 citations · 1 across the 9 of their papers we have counts for
7 papers · 1 filter
FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning
Haihui Pan, Junwei Bao, Hongfei Jiang +1
While large language models have made significant progress in mathematical reasoning, they remain unreliable at judging the correctness of their own solutions. Existing approaches…
Elo-Evolve: A Co-evolutionary Framework for Language Model Alignment
Jing Zhao, Ting Zhen, Junwei Bao +2
Current alignment methods for Large Language Models (LLMs) rely on compressing vast amounts of human preference data into static, absolute reward functions, leading to data scarcit…
Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity
Haihui Pan, Yuzhong Hong, Kaichen Zhang +4
In many large language model (LLM) alignment applications, users expect not only high-quality outputs but also substantial diversity. However, existing methods often face a fundame…
Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models
Yuchen Fan, Yuzhong Hong, Qiushi Wang +3
Alignment, endowing a pre-trained Large language model (LLM) with the ability to follow instructions, is crucial for its real-world applications. Conventional supervised fine-tunin…
BoRA: Bi-dimensional Weight-Decomposed Low-Rank Adaptation
Qiushi Wang, Yuchen Fan, Junwei Bao +2
In recent years, Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA) have significantly enhanced the adaptability of large-scale pre-trained models. Weig…
Multi-Turn Interactions for Text-to-SQL with Large Language Models
Guanming Xiong, Junwei Bao, Hongfei Jiang +2
This study explores text-to-SQL parsing by leveraging the powerful reasoning capabilities of large language models (LLMs). Despite recent advancements, existing LLM-based methods a…