collaborators

8 papers

cs.AI2026

When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction

Feiyang Ren, Shengtao Wen, Lingbing Guo +3

Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades throughput. Predicting output lengths in ad…

cs.AI2026

VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation

Mingyu Yuan, Shengtao Wen, Lingbing Guo +2

The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing Chinese benchmarks support label classifi…

cs.AI2026

Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing

Sheng Ren, Yadong Wang, Naiqiang Tan +7

Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when…

stat.ML2026

Beyond Importance: Interchange-Sobol Sensitivity Reveals Task-Specific Content Channels in Transformer Components

Yifeng Guo, Jin-Hong Du, Xiang Chen

Mechanistic interpretability methods summarize a transformer component by a single importance score, conflating two distinct roles: a component may matter because it transports tas…

cs.CV2026

CoDA: Exploring Chain-of-Distribution Attacks and Post-Hoc Token-Space Repair for Medical Vision-Language Models

Xiang Chen, Fangfang Yang, Chunlei Meng +6

Medical vision--language models (MVLMs) are increasingly used as perceptual backbones in radiology pipelines and as the visual front end of multimodal assistants, yet their reliabi…

cs.CR2026

SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment

Xianya Fang, Xianying Luo, Yadong Wang +8

Despite the intrinsic risk-awareness of Large Language Models (LLMs), current defenses often result in shallow safety alignment, rendering models vulnerable to disguised attacks (e…