10 papers
Smooth Scaling Laws Hide Stepwise Token Learning
Pingjie Wang, Zechen Hu, Peiru Yang +2
Language model loss follows remarkably regular scaling laws over model and data size, yet it remains unclear why the aggregate loss should exhibit a power-law form. Existing explan…
Knowledge Poisoning Attacks on Medical Multi-Modal Retrieval-Augmented Generation
Peiru Yang, Haoran Zheng, Tong Ju +6
Retrieval-augmented generation (RAG) is a widely adopted paradigm for enhancing LLMs in medical applications by incorporating expert multimodal knowledge during generation. However…
RedNote-Vibe: A Dataset for Capturing Temporal Dynamics of AI-Generated Text in Lifestyle Social Media
Yudong Li, Yufei Sun, Peiru Yang +7
We introduce RedNote-Vibe, a dataset spanning five years (pre-LLM to July 2025) sourced from lifestyle platform RedNote (Xiaohongshu), capturing the temporal dynamics of content cr…
Igusa-Todorov properties of recollements of abelian categories
Peiru Yang, Yajun Ma, Yu-Zhe Liu
In this paper, we investigate the behavior of Igusa-Todorov properties under recollements of abelian categories. In particular, we study how the Igusa-Todorov distances of the cate…
LiveSecBench: A Dynamic and Event-Driven Safety Benchmark for Chinese Language Model Applications
Yudong Li, Peiru Yang, Feng Huang +18
We introduce LiveSecBench, a continuously updated safety benchmark specifically for Chinese-language LLM application scenarios. LiveSecBench constructs a high-quality and unique da…
Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing
Jinhua Yin, Peiru Yang, Chen Yang +5
Large vision-language models (LVLMs) derive their capabilities from extensive training on vast corpora of visual and textual data. Empowered by large-scale parameters, these models…