activity
20242026
most citedWhen Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models

1 citations · 1 across the 7 of their papers we have counts for

collaborators

9 papers

cs.LG2026

Absorber LLM: Harnessing Causal Synchronization for Test-Time Training

Zhixin Zhang, Shabo Zhang, Chengcan Wu +2

Transformers suffer from a high computational cost that grows with sequence length for self-attention, making inference in long streams prohibited by memory consumption. Constant-m…

cs.CR2026

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems

Yihao Zhang, Kai Wang, Jiangrong Wu +7

Large Language Models (LLMs) face prominent security risks from jailbreaking, a practice that manipulates models to bypass built-in security constraints and generate unethical or u…

cs.CL2025

Automata-Based Steering of Large Language Models for Diverse Structured Generation

Xiaokun Luan, Zeming Wei, Yihao Zhang +1

Large language models (LLMs) are increasingly tasked with generating structured outputs. While structured generation methods ensure validity, they often lack output diversity, a cr…

cs.LG2025

Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings

Zhixin Zhang, Zeming Wei, Meng Sun

Catastrophic forgetting remains a critical challenge in continual learning for large language models (LLMs), where models struggle to retain performance on historical tasks when fi…

cs.LG2025

Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection

Chengcan Wu, Zeming Wei, Huanran Chen +2

While Large Language Models (LLMs) have demonstrated impressive performance in various domains and tasks, concerns about their safety are becoming increasingly severe. In particula…

cs.AI20251 cited

When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models

Kai Wang, Yihao Zhang, Meng Sun

The honesty of large language models (LLMs) is a critical alignment challenge, especially as advanced systems with chain-of-thought (CoT) reasoning may strategically deceive humans…