1 citations · 1 across the 12 of their papers we have counts for
7 papers · 1 filter
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
Binghai Wang, Yantao Liu, Yuxuan Liu +13
Generative Reward Models (GenRMs) and LLM-as-a-Judge exhibit deceptive alignment by producing correct judgments for incorrect reasons, as they are trained and evaluated to prioriti…
RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning
Xiang Gao, Yuguang Yao, Qi Zhang +5
Large language models (LLMs) often struggle to use tools reliably in domain-specific settings, where APIs may be idiosyncratic, under-documented, or tailored to private workflows.…
ChemATP: A Training-Free Chemical Reasoning Framework for Large Language Models
Mingxu Zhang, Dazhong Shen, Qi Zhang +1
Large Language Models (LLMs) exhibit strong general reasoning but struggle in molecular science due to the lack of explicit chemical priors in standard string representations. Curr…
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Chong Zhang, Yue Deng, Xiang Lin +8
The recent development of reasoning language models (RLMs) represents a novel evolution in large language models. In particular, the recent release of DeepSeek-R1 has generated wid…
A Comprehensive Framework for Semantic Similarity Analysis of Human and AI-Generated Text Using Transformer Architectures and Ensemble Techniques
Lifu Gao, Ziwei Liu, Qi Zhang
The rapid advancement of large language models (LLMs) has made detecting AI-generated text an increasingly critical challenge. Traditional methods often fail to capture the nuanced…
Optimizing Sentence Embedding with Pseudo-Labeling and Model Ensembles: A Hierarchical Framework for Enhanced NLP Tasks
Ziwei Liu, Qi Zhang, Lifu Gao
Sentence embedding tasks are important in natural language processing (NLP), but improving their performance while keeping them reliable is still hard. This paper presents a framew…