activity
20242026
collaborators

65 papers

cs.CL2026

ZenGen: Social Mind for LLMs

ZenGen Team, Zing Team, Ao Xiang +57

As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track…

cs.CR2026

Token-Flow Firewall: Semantic Runtime Auditing for Persistent AI Agents

Puji Wang, Yingchen Zhang, Ruqing Zhang +2

Persistent AI agents extend large language models (LLMs) beyond single-turn interaction into long-lived software systems. Unlike traditional chat assistants, unsafe content in thes…

cs.IR2026

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning

Jiahan Chen, Da Li, Hengran Zhang +6

Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classi…

cs.CL2026

EGAD: Entropy-Guided Adaptive Distillation for Token-Level Knowledge Transfer

Hao Zhang, Zhibin Zhang, Guangxin Wu +3

Large language models (LLMs) have achieved remarkable performance across diverse domains, yet their enormous computational and memory requirements hinder deployment in resource-con…

cs.CL2026

Detoxification for LLM: From Dataset Itself

Wei Shao, Yihang Wang, Gaoyu Zhu +4

Existing detoxification methods for large language models mainly focus on post-training stage or inference time, while few tackle the source of toxicity, namely, the dataset itself…

cs.CV2026

Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding

Da Li, Yuxiao Luo, Keping Bi +7

Multimodal Large Language Models advance multimodal representation learning by acquiring transferable semantic embeddings, thereby substantially enhancing performance across a rang…