8 papers
When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction
Feiyang Ren, Shengtao Wen, Lingbing Guo +3
Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades throughput. Predicting output lengths in ad…
VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation
Mingyu Yuan, Shengtao Wen, Lingbing Guo +2
The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing Chinese benchmarks support label classifi…
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing
Sheng Ren, Yadong Wang, Naiqiang Tan +7
Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when…
Beyond Importance: Interchange-Sobol Sensitivity Reveals Task-Specific Content Channels in Transformer Components
Yifeng Guo, Jin-Hong Du, Xiang Chen
Mechanistic interpretability methods summarize a transformer component by a single importance score, conflating two distinct roles: a component may matter because it transports tas…
CoDA: Exploring Chain-of-Distribution Attacks and Post-Hoc Token-Space Repair for Medical Vision-Language Models
Xiang Chen, Fangfang Yang, Chunlei Meng +6
Medical vision--language models (MVLMs) are increasingly used as perceptual backbones in radiology pipelines and as the visual front end of multimodal assistants, yet their reliabi…
SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment
Xianya Fang, Xianying Luo, Yadong Wang +8
Despite the intrinsic risk-awareness of Large Language Models (LLMs), current defenses often result in shallow safety alignment, rendering models vulnerable to disguised attacks (e…