5 papers
The Bidirectional Process Reward Model
Lingyin Zhang, Jun Gao, Xiaoxue Ren +1
Process Reward Models (PRMs), which assign fine-grained scores to intermediate reasoning steps within a solution trajectory, have emerged as a promising approach to enhance the rea…
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
Qi Lv, Lei Geng, Ziqiang Cao +4
Softmax with the cross entropy loss is the standard configuration for current neural classification models. The gold score for a target class is supposed to be 1, but it is never r…
UniICL: An Efficient Unified Framework Unifying Compression, Selection, and Generation
Jun Gao, Qi Lv, Zili Wang +3
In-context learning (ICL) enhances the reasoning abilities of Large Language Models (LLMs) by prepending a few demonstrations. It motivates researchers to introduce more examples t…
Interleaved-Modal Chain-of-Thought
Jun Gao, Yongqi Li, Ziqiang Cao +1
Chain-of-Thought (CoT) prompting elicits large language models (LLMs) to produce a series of intermediate reasoning steps before arriving at the final answer. However, when transit…
Personalized Large Language Model Assistant with Evolving Conditional Memory
Ruifeng Yuan, Shichao Sun, Yongqi Li +3
With the rapid development of large language models, AI assistants like ChatGPT have become increasingly integrated into people's works and lives but are limited in personalized se…