6 papers
A Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-controlled and Dynamic Test Generation
Shide Zhou, Kailong Wang, Ling Shi +1
Defining the reasoning boundaries and ensuring the reliability of Large Reasoning Models (LRMs) remains a critical challenge. Current benchmarks primarily rely on static datasets s…
Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics
Shide Zhou, Kailong Wang, Ling Shi +1
The widespread adoption of Large Language Models (LLMs) in critical applications has introduced severe reliability and security risks, as LLMs remain vulnerable to notorious threat…
When Safe Models Merge into Danger: Exploiting Latent Vulnerabilities in LLM Fusion
Jiaqing Li, Zhibo Zhang, Shide Zhou +3
Model merging has emerged as a powerful technique for combining specialized capabilities from multiple fine-tuned LLMs without additional training costs. However, the security impl…
OmniBench-RAG: A Multi-Domain Evaluation Platform for Retrieval-Augmented Generation Tools
Jiaxuan Liang, Shide Zhou, Kailong Wang
While Retrieval Augmented Generation (RAG) is now widely adopted to enhance LLMs, evaluating its true performance benefits in a reproducible and interpretable way remains a major h…
NeuSemSlice: Towards Effective DNN Model Maintenance via Neuron-level Semantic Slicing
Shide Zhou, Tianlin Li, Yihao Huang +4
Deep Neural networks (DNNs), extensively applied across diverse disciplines, are characterized by their integrated and monolithic architectures, setting them apart from conventiona…
Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks
Shide Zhou, Tianlin Li, Kailong Wang +4
Large language models (LLMs) have revolutionized artificial intelligence, but their increasing deployment across critical domains has raised concerns about their abnormal behaviors…