collaborators

5 papers

cs.CL2025

A Survey on Large Language Model Benchmarks

Shiwen Ni, Guhong Chen, Shuaimin Li +11

In recent years, with the rapid development of the depth and breadth of large language models' capabilities, various corresponding evaluation benchmarks have been emerging in incre…

cs.CL2025

Training on the Benchmark Is Not All You Need

Shiwen Ni, Xiangtao Kong, Chengming Li +4

The success of Large Language Models (LLMs) relies heavily on the huge amount of pre-training data learned in the pre-training phase. The opacity of the pre-training process and th…

cs.CL2025

xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking

Sunbowen Lee, Shiwen Ni, Chi Wei +7

Safety alignment mechanism are essential for preventing large language models (LLMs) from generating harmful information or unethical content. However, cleverly crafted prompts can…

cs.CL2025

AutoCBT: An Autonomous Multi-agent Framework for Cognitive Behavioral Therapy in Psychological Counseling

Ancheng Xu, Di Yang, Renhao Li +13

Traditional in-person psychological counseling remains primarily niche, often chosen by individuals with psychological issues, while online automated counseling offers a potential…

cs.CL2025

II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models

Ziqiang Liu, Feiteng Fang, Xi Feng +23

The rapid advancements in the development of multimodal large language models (MLLMs) have consistently led to new breakthroughs on various benchmarks. In response, numerous challe…