most citedInsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models

1 citations · 2 across the 4 of their papers we have counts for

collaborators

6 papers

cs.LG2025

TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs

Yuxiang Zhang, Zhengxu Yu, Weihang Pan +5

Emerging reasoning LLMs such as OpenAI-o1 and DeepSeek-R1 have achieved strong performance on complex reasoning tasks by generating long chain-of-thought (CoT) traces. However, the…

cs.CL2025

Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning

Chenxi Huang, Shaotian Yan, Liang Xie +6

Representation Fine-tuning (ReFT), a recently proposed Parameter-Efficient Fine-Tuning (PEFT) method, has attracted widespread attention for significantly improving parameter effic…

cs.LG20251 cited

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning

Mengsong Wu, YaFei Wang, Yidong Ming +7

Large language models (LLMs) have recently demonstrated promising capabilities in chemistry tasks while still facing challenges due to outdated pretraining knowledge and the diffic…

cs.CL20251 cited

InsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models

Jing Ding, Kai Feng, Binbin Lin +6

The application of large language models (LLMs) has achieved remarkable success in various fields, but their effectiveness in specialized domains like the Chinese insurance industr…

cs.CL2024

Delving into the Reversal Curse: How Far Can Large Language Models Generalize?

Zhengkai Lin, Zhihang Fu, Kai Liu +6

While large language models (LLMs) showcase unprecedented capabilities, they also exhibit certain inherent limitations when facing seemingly trivial tasks. A prime example is the r…

cs.CL2024

Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control

Yuxin Xiao, Chaoqun Wan, Yonggang Zhang +5

As the development and application of Large Language Models (LLMs) continue to advance rapidly, enhancing their trustworthiness and aligning them with human preferences has become…