papers

Publications (45)

cs.CL2026

Learning to Edit Knowledge via Instruction-based Chain-of-Thought Prompting

Jinhu Fu, Yan Bai, Longzhu He +4

Large language models (LLMs) can effectively handle outdated information through knowledge editing. However, current approaches face two key limitations: (I) Poor generalization: M…

cs.CV2021

Dense Contrastive Visual-Linguistic Pretraining

Lei Shi, Kai Shuang, Shijie Geng +5

Inspired by the success of BERT, several multimodal representation learning approaches have been proposed that jointly represent image and text. These approaches achieve superior p…

cs.CL2024

Alignment-Enhanced Decoding:Defending via Token-Level Adaptive Refining of Probability Distributions

Quan Liu, Zhenhong Zhou, Longzhu He +3

Large language models are susceptible to jailbreak attacks, which can result in the generation of harmful content. While prior defenses mitigate these risks by perturbing or inspec…

cs.CL2024

Smaller Language Models Are Better Instruction Evolvers

Tingfeng Hui, Lulu Zhao, Guanting Dong +3

Instruction tuning has been widely used to unleash the complete potential of large language models. Notably, complex and diverse instructions are of significant importance as they…

cs.CV2020

Contrastive Visual-Linguistic Pretraining

Lei Shi, Kai Shuang, Shijie Geng +6

Several multi-modality representation learning approaches such as LXMERT and ViLBERT have been proposed recently. Such approaches can achieve superior performance due to the high-l…

cs.AI2026

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

Pengyu Zhu, Lijun Li, Longju Yang +2

Deep Research agents conduct long-horizon investigations by iteratively planning, retrieving evidence, and generating reports. However, it remains unclear whether they can resist a…

cs.LG2026

Towards Personalized Differentially Private Learning for Decentralized Local Graphs

Longzhu He, Peng Tang, Chaozhuo Li +5

Graph-structured data is increasingly generated and stored in decentralized environments, such as social platforms, mobile applications, and edge networks, where users maintain con…

cs.CL2023

Quantifying and Analyzing Entity-level Memorization in Large Language Models

Zhenhong Zhou, Jiuyang Xiang, Chaomeng Chen +1

Large language models (LLMs) have been proven capable of memorizing their training data, which can be extracted through specifically designed prompts. As the scale of datasets cont…

cs.CR2026

Resource Consumption Threats in Large Language Models

Yuanhe Zhang, Xinyue Wang, Zhican Chen +8

Given limited and costly computational infrastructure, resource efficiency is a key requirement for large language models (LLMs). Efficient LLMs increase service capacity for provi…

cs.CV2020

Multi-Layer Content Interaction Through Quaternion Product For Visual Question Answering

Lei Shi, Shijie Geng, Kai Shuang +4

Multi-modality fusion technologies have greatly improved the performance of neural network-based Video Description/Caption, Visual Question Answering (VQA) and Audio Visual Scene-a…

cs.SD2026

SEE: Signal Embedding Energy for Quantifying Noise Interference in Large Audio Language Models

Yuanhe Zhang, Jiayu Tian, Yibo Zhang +5

Large Audio Language Models (LALMs) have been widely applied in real-time scenarios, such as in-car assistants and online meeting comprehension. In practice, audio inputs are often…

cs.CV2024

Unsupervised Attention Regularization Based Domain Adaptation for Oracle Character Recognition

Mei Wang, Weihong Deng, Jiani Hu +1

The study of oracle characters plays an important role in Chinese archaeology and philology. However, the difficulty of collecting and annotating real-world scanned oracle characte…

cs.CL2021

Adaptive Noise Injection: A Structure-Expanding Regularization for RNN

Rui Li, Kai Shuang, Mengyu Gu +1

The vanilla LSTM has become one of the most potential architectures in word-level language modeling, like other recurrent neural networks, overfitting is always a key barrier for i…

cs.CV2025

Mitigating Group-Level Fairness Disparities in Federated Visual Language Models

Chaomeng Chen, Zitong Yu, Junhao Dong +4

Visual language models (VLMs) have shown remarkable capabilities in multimodal tasks but face challenges in maintaining fairness across demographic groups, particularly when deploy…

cs.CL2023

BitCoin: Bidirectional Tagging and Supervised Contrastive Learning based Joint Relational Triple Extraction Framework

Luyao He, Zhongbao Zhang, Sen Su +1

Relation triple extraction (RTE) is an essential task in information extraction and knowledge graph construction. Despite recent advancements, existing methods still exhibit certai…

cs.AI2026

"LLM Agent Performance" Is Not a Single Evaluation Target

Pengyu Zhu, Li Sun, Philip S. Yu +1

LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget. Unified execution controls these non-model…

cs.AI2026

A Unified Framework for the Evaluation of LLM Agentic Capabilities

Pengyu Zhu, Lijun Li, Yaxing Lyu +8

As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores often jointly reflect model…

cs.CV2026

Structure-Guided Visual Perturbation Neutralization for LVLMs

Yuanhe Zhang, Xueting Wang, YanBin Ren +6

Image inputs enable Large Vision Language Models (LVLMs) to perceive fine-grained visual information, but also introduce a pixel-level attack surface through which adversarial pert…

cs.CR2025

Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems

Pengyu Zhu, Lijun Li, Yaxing Lyu +3

LLM-based multi-agent systems (MAS) demonstrate increasing integration into next-generation applications, but their safety in backdoor attacks remains largely underexplored. Howeve…

cs.AI2026

Heterophily-Agnostic Hypergraph Neural Networks with Riemannian Local Exchanger

Li Sun, Ming Zhang, Wenxin Jin +5

Hypergraphs are the natural description of higher-order interactions among objects, widely applied in social network analysis, cross-modal retrieval, etc. Hypergraph Neural Network…

cs.CL2025

Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings

Yuanhe Zhang, Zhenhong Zhou, Wei Zhang +4

Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks yet still are vulnerable to external threats, particularly LLM Denial-of-Service (LLM-DoS…

cs.CR2025

: A Pluggable and Dynamic DoS-Defense Framework Against Resource Consumption Attacks Targeting Large Language Models

Yuanhe Zhang, Xinyue Wang, Haoran Gao +4

Large Language Models (LLMs), due to substantial computational requirements, are vulnerable to resource consumption attacks, which can severely degrade server performance or even c…

cs.CR2020

Answering Multi-Dimensional Range Queries under Local Differential Privacy

Jianyu Yang, Tianhao Wang, Ninghui Li +2

In this paper, we tackle the problem of answering multi-dimensional range queries under local differential privacy. There are three key technical challenges: capturing the correlat…

cs.CL2025

DecIF: Improving Instruction-Following through Meta-Decomposition

Tingfeng Hui, Pengyu Zhu, Bowen Ping +4

Instruction-following has emerged as a crucial capability for large language models (LLMs). However, existing approaches often rely on pre-existing documents or external resources…

cs.CL2026

From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents

Xinyue Wang, Yuanhe Zhang, Zhengshuo Gong +6

The enhanced capabilities of LLM-based agents come with an emergency for model planning and tool-use abilities. Attributing to helpful-harmless trade-off from LLM alignment, agents…

cs.LG2021

Hyperbolic Variational Graph Neural Network for Modeling Dynamic Graphs

Li Sun, Zhongbao Zhang, Jiawei Zhang +4

Learning representations for graphs plays a critical role in a wide spectrum of downstream applications. In this paper, we summarize the limitations of the prior works in three fol…

cs.CV2023

Oracle Character Recognition using Unsupervised Discriminative Consistency Network

Mei Wang, Weihong Deng, Sen Su

Ancient history relies on the study of ancient characters. However, real-world scanned oracle characters are difficult to collect and annotate, posing a major obstacle for oracle c…

cs.CL2026

LARFT: Closing the Cognition-Action Gap for Length Instruction Following in Large Language Models

Wei Zhang, Lintong Du, Yuanhe Zhang +4

Despite the strong performance of Large Language Models (LLMs) on complex instruction-following tasks, precise control of output length remains a persistent challenge. Existing met…

cs.CL2024

Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging

Tingfeng Hui, Zhenyu Zhang, Shuohuan Wang +3

Mixture-of-Experts (MoE) shines brightly in large language models (LLMs) and demonstrates outstanding performance in plentiful natural language processing tasks. However, existing…

cs.LG2025

Devil's Hand: Data Poisoning Attacks to Locally Private Graph Learning Protocols

Longzhu He, Chaozhuo Li, Peng Tang +3

Graph neural networks (GNNs) have achieved significant success in graph representation learning and have been applied to various domains. However, many real-world graphs contain se…

cs.CV2024

Marginal Debiased Network for Fair Visual Recognition

Mei Wang, Weihong Deng, Jiani Hu +1

Deep neural networks (DNNs) are often prone to learn the spurious correlations between target classes and bias attributes, like gender and race, inherent in a major portion of trai…

cs.CR2025

LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems

Yuanhe Zhang, Weiliu Wang, Zhenhong Zhou +5

Large Language Model (LLM)-based agents have demonstrated remarkable capabilities in reasoning, planning, and tool usage. The recently proposed Model Context Protocol (MCP) has eme…

cs.AI2026

From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control

Rui Ha, Rui Pu, Chaozhuo Li +2

Large Reasoning Models (LRMs) can exhibit step-by-step reasoning, reflection, and backtracking, but these behaviors are often unregulated, leading to overthinking. As a result, LRM…

cs.AI2026

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation

Yi Liu, TingFeng Hui, Wei Zhang +4

Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive to build, bri…

cs.SI2022

PERFECT: A Hyperbolic Embedding for Joint User and Community Alignment

Li Sun, Zhongbao Zhang, Jiawei Zhang +4

Social network alignment shows fundamental importance in a wide spectrum of applications. To the best of our knowledge, existing studies mainly focus on network alignment at the in…

cs.CV2026

Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection

Jinhu Fu, Yihang Lou, Qingyi Si +3

Large Vision-Language Models (LVLMs) have achieved impressive performance across multimodal understanding and reasoning tasks, yet their internal safety mechanisms remain opaque an…

cs.CR2025

DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent

Pengyu Zhu, Zhenhong Zhou, Yuanhe Zhang +3

As LLM-based agents become increasingly prevalent, backdoors can be implanted into agents through user queries or environment feedback, raising critical concerns regarding safety v…

cs.CL2026

STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics

Tingfeng Hui, Hao Xu, Pengyu Zhu +5

Large language models (LLMs) deployed in real-world agentic applications must be capable of replanning and adapting when mid-task disruptions invalidate their prior decisions. Exis…

cs.CL2025

LIFEBench: Evaluating Length Instruction Following in Large Language Models

Wei Zhang, Zhenhong Zhou, Kun Wang +9

While large language models (LLMs) can solve PhD-level reasoning problems over long context inputs, they still struggle with a seemingly simpler task: following explicit length ins…

cs.LG2021

A Self-supervised Mixed-curvature Graph Neural Network

Li Sun, Zhongbao Zhang, Junda Ye +4

Graph representation learning received increasing attentions in recent years. Most of existing methods ignore the complexity of the graph structures and restrict graphs in a single…

cs.LG2026

Multi-Domain Riemannian Graph Gluing for Building Graph Foundation Models

Li Sun, Zhenhao Huang, Silei Chen +4

Multi-domain graph pre-training integrates knowledge from diverse domains to enhance performance in the target domains, which is crucial for building graph foundation models. Despi…

cs.CL2024

Speak Out of Turn: Safety Vulnerability of Large Language Models in Multi-turn Dialogue

Zhenhong Zhou, Jiuyang Xiang, Haopeng Chen +3

Large Language Models (LLMs) have been demonstrated to generate illegal or unethical responses, particularly when subjected to "jailbreak." Research on jailbreak has highlighted th…

cs.SI2019

DNA: Dynamic Social Network Alignment

Li Sun, Zhongbao Zhang, Pengxin Ji +3

Social network alignment, aligning different social networks on their common users, is receiving dramatic attention from both academic and industry. All existing studies consider t…

cs.LG2026

Residual Stream Analysis of Overfitting And Structural Disruptions

Quan Liu, Han Zhou, Wenquan Wu +2

Ensuring that large language models (LLMs) remain both helpful and harmless poses a significant challenge: fine-tuning on repetitive safety datasets, where unsafe prompts are paire…

cs.SD2026

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

Kaiwen Luo, Zhenhong Zhou, Leo Wang +34

Advances in Large Language Models (LLMs) have paved the way for Multimodal Large Language Models (MLLMs). Among these, Large Audio Language Models (LALMs) are essential for realizi…