papers

Publications (274)

cs.CR2025

Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation

Yinuo Liu, Zenghui Yuan, Guiyao Tie +4

Multimodal retrieval-augmented generation (RAG) enhances the visual reasoning capability of vision-language models (VLMs) by dynamically accessing information from external knowled…

cs.LG2019

Exact Recovery of Tensor Robust Principal Component Analysis under Linear Transforms

Canyi Lu, Pan Zhou

This work studies the Tensor Robust Principal Component Analysis (TRPCA) problem, which aims to exactly recover the low-rank and sparse components from their sum. Our model is moti…

cs.CY2018

DeepSpace: An Online Deep Learning Framework for Mobile Big Data to Understand Human Mobility Patterns

Xi Ouyang, Chaoyun Zhang, Pan Zhou +2

In the recent years, the rapid spread of mobile device has create the vast amount of mobile data. However, some shallow-structure models such as support vector machine (SVM) have d…

cs.CR2025

Prompt Injection Attack to Tool Selection in LLM Agents

Jiawen Shi, Zenghui Yuan, Guiyao Tie +3

Tool selection is a key component of LLM agents. A popular approach follows a two-step process - \emph{retrieval} and \emph{selection} - to pick the most appropriate tool from a to…

cs.CV2020

F2Net: Learning to Focus on the Foreground for Unsupervised Video Object Segmentation

Daizong Liu, Dongdong Yu, Changhu Wang +1

Although deep learning based methods have achieved great progress in unsupervised video object segmentation, difficult scenarios (e.g., visual similarity, occlusions, and appearanc…

cs.CV2024

Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator

Henry Hengyuan Zhao, Pan Zhou, Mike Zheng Shou

Multimodal Large Language Models (MLLMs) demonstrate exceptional problem-solving capabilities, but few research studies aim to gauge the ability to generate visual instruction tuni…

cs.LG2024

Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models

Xingyu Xie, Pan Zhou, Huan Li +2

In deep learning, different kinds of deep networks typically need different optimizers, which have to be chosen after multiple trials, making the training process inefficient. To r…

cs.CV2025

Towards Scalable and Consistent 3D Editing

Ruihao Xia, Yang Tang, Pan Zhou

3D editing - the task of locally modifying the geometry or appearance of a 3D asset - has wide applications in immersive content creation, digital entertainment, and AR/VR. However…

cs.LG2020

Theory-Inspired Path-Regularized Differential Network Architecture Search

Pan Zhou, Caiming Xiong, Richard Socher +1

Despite its high search efficiency, differential architecture search (DARTS) often selects network architectures with dominated skip connections which lead to performance degradati…

cs.CV2020

Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment Localization

Daizong Liu, Xiaoye Qu, Xiao-Yang Liu +3

Query-based moment localization is a new task that localizes the best matched segment in an untrimmed video according to a given sentence query. In this localization task, one shou…

cs.CL2024

Way to Specialist: Closing Loop Between Specialized LLM and Evolving Domain Knowledge Graph

Yutong Zhang, Lixing Chen, Shenghong Li +6

Large language models (LLMs) have demonstrated exceptional performance across a wide variety of domains. Nonetheless, generalist LLMs continue to fall short in reasoning tasks nece…

cs.CR2024

Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection

Yuqi Zhou, Lin Lu, Hanchi Sun +2

Jailbreak attacks on large language models (LLMs) involve inducing these models to generate harmful content that violates ethics or laws, posing a significant threat to LLM securit…

cs.CV2022

Progressive Localization Networks for Language-based Moment Localization

Qi Zheng, Jianfeng Dong, Xiaoye Qu +5

This paper targets the task of language-based video moment localization. The language-based setting of this task allows for an open set of target activities, resulting in a large v…

cs.RO2025

Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents

Zhejian Yang, Yongchao Chen, Xueyang Zhou +8

Long-horizon robotic manipulation poses significant challenges for autonomous systems, requiring extended reasoning, precise execution, and robust error recovery across complex seq…

cs.CV2022

MetaFormer Is Actually What You Need for Vision

Weihao Yu, Mi Luo, Pan Zhou +5

Transformers have shown great potential in computer vision tasks. A common belief is their attention-based token mixer module contributes most to their competence. However, recent…

cs.LG2025

4-bit Shampoo for Memory-Efficient Network Training

Sike Wang, Pan Zhou, Jia Li +1

Second-order optimizers, maintaining a matrix termed a preconditioner, are superior to first-order optimizers in both theory and practice. The states forming the preconditioner and…

cs.CV2024

Instant3D: Instant Text-to-3D Generation

Ming Li, Pan Zhou, Jia-Wei Liu +4

Text-to-3D generation has attracted much attention from the computer vision community. Existing methods mainly optimize a neural field from scratch for each text prompt, relying on…

cs.CV2025

ReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment

Wanjiang Weng, Xiaofeng Tan, Hongsong Wang +1

Bilingual text-to-motion generation, which synthesizes 3D human motions from bilingual text inputs, holds immense potential for cross-linguistic applications in gaming, film, and r…

cs.CR2023

GraphCloak: Safeguarding Task-specific Knowledge within Graph-structured Data from Unauthorized Exploitation

Yixin Liu, Chenrui Fan, Xun Chen +2

As Graph Neural Networks (GNNs) become increasingly prevalent in a variety of fields, from social network analysis to protein-protein interaction studies, growing concerns have eme…

cs.CV2022

Self-Promoted Supervision for Few-Shot Transformer

Bowen Dong, Pan Zhou, Shuicheng Yan +1

The few-shot learning ability of vision transformers (ViTs) is rarely investigated though heavily desired. In this work, we empirically find that with the same few-shot learning fr…

cs.SD2023

Automatic channel selection and spatial feature integration for multi-channel speech recognition across various array topologies

Bingshen Mu, Pengcheng Guo, Dake Guo +3

Automatic Speech Recognition (ASR) has shown remarkable progress, yet it still faces challenges in real-world distant scenarios across various array topologies each with multiple r…

cs.CR2023

BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Jiawen Shi, Yixin Liu, Pan Zhou +1

Recently, ChatGPT has gained significant attention in research due to its ability to interact with humans effectively. The core idea behind this model is reinforcement learning (RL…

cs.DC2020

Communication-efficient Decentralized Machine Learning over Heterogeneous Networks

Pan Zhou, Qian Lin, Dumitrel Loghin +3

In the last few years, distributed machine learning has been usually executed over heterogeneous networks such as a local area network within a multi-tenant cluster or a wide area…

eess.AS2023

U2-KWS: Unified Two-pass Open-vocabulary Keyword Spotting with Keyword Bias

Ao Zhang, Pan Zhou, Kaixun Huang +3

Open-vocabulary keyword spotting (KWS), which allows users to customize keywords, has attracted increasingly more interest. However, existing methods based on acoustic models and p…

cs.LG2015

Differentially Private Distributed Online Learning

Chencheng Li, Pan Zhou

Online learning has been in the spotlight from the machine learning society for a long time. To handle massive data in Big Data era, one single learner could never efficiently fini…

cs.CV2024

What Makes Good Collaborative Views? Contrastive Mutual Information Maximization for Multi-Agent Perception

Wanfang Su, Lixing Chen, Yang Bai +4

Multi-agent perception (MAP) allows autonomous systems to understand complex environments by interpreting data from multiple sources. This paper investigates intermediate collabora…

cs.CL2019

Exploring RNN-Transducer for Chinese Speech Recognition

Senmao Wang, Pan Zhou, Wei Chen +2

End-to-end approaches have drawn much attention recently for significantly simplifying the construction of an automatic speech recognition (ASR) system. RNN transducer (RNN-T) is o…

cond-mat.mtrl-sci2026

High-Throughput Discovery of Two-Dimensional Materials Exhibiting Strong Rashba-Edelstein effect

Binchang Zhou, Baoru Pan, Pan Zhou +3

The Rashba-Edelstein effect (REE), which generates spin accumulation under an applied electric current, quantifies charge-to-spin conversion (CSC) efficiency in non-centrosymmetric…

cs.CL2025

GRIFFIN: Effective Token Alignment for Faster Speculative Decoding

Shijing Hu, Jingyang Li, Xingyu Xie +3

Speculative decoding accelerates inference in large language models (LLMs) by generating multiple draft tokens simultaneously. However, existing methods often struggle with token m…

cs.CV2024

A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Daizong Liu, Mingyu Yang, Xiaoye Qu +3

With the significant development of large models in recent years, Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal u…

cs.AI2020

Transfer Heterogeneous Knowledge Among Peer-to-Peer Teammates: A Model Distillation Approach

Zeyue Xue, Shuang Luo, Chao Wu +3

Peer-to-peer knowledge transfer in distributed environments has emerged as a promising method since it could accelerate learning and improve team-wide performance without relying o…

eess.AS2024

The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023

He Wang, Pengcheng Guo, Wei Chen +2

This paper delineates the visual speech recognition (VSR) system introduced by the NPU-ASLP-LiAuto (Team 237) in the first Chinese Continuous Visual Speech Recognition Challenge (C…

cs.LG2026

IBNorm: Information-Bottleneck Inspired Normalization for Representation Learning

Xiandong Zou, Jia Li, Xiaotong Yuan +1

Normalization is fundamental to deep learning, but existing approaches such as BatchNorm, LayerNorm, and RMSNorm are variance-centric by enforcing zero mean and unit variance, stab…

cs.CV2025

Temporal-Guided Visual Foundation Models for Event-Based Vision

Ruihao Xia, Junhong Cai, Luziwei Leng +5

Event cameras offer unique advantages for vision tasks in challenging environments, yet processing asynchronous event streams remains an open challenge. While existing methods rely…

cs.CV2026

LiAuto-GeoX: Efficient Grounded Driving Transformer

Jiawei Lian, Haoyi Sun, Yang Wu +8

Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for autonomous driving remains an ope…

cs.LG2024

Friendly Sharpness-Aware Minimization

Tao Li, Pan Zhou, Zhengbao He +2

Sharpness-Aware Minimization (SAM) has been instrumental in improving deep neural network training by minimizing both training loss and loss sharpness. Despite the practical succes…

cs.CV2023

STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition

Ming Li, Xiangyu Xu, Hehe Fan +7

Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks…

cs.CV2021

A Theory-Driven Self-Labeling Refinement Method for Contrastive Representation Learning

Pan Zhou, Caiming Xiong, Xiao-Tong Yuan +1

For an image query, unsupervised contrastive learning labels crops of the same image as positives, and other image crops as negatives. Although intuitive, such a native label assig…

cs.CV2023

3DHacker: Spectrum-based Decision Boundary Generation for Hard-label 3D Point Cloud Attack

Yunbo Tao, Daizong Liu, Pan Zhou +3

With the maturity of depth sensors, the vulnerability of 3D point cloud models has received increasing attention in various applications such as autonomous driving and robot naviga…

cs.CR2023

Backdoor Attacks to Pre-trained Unified Foundation Models

Zenghui Yuan, Yixin Liu, Kai Zhang +2

The rise of pre-trained unified foundation models breaks down the barriers between different modalities and tasks, providing comprehensive support to users with unified architectur…

cs.LG2021

How Important is the Train-Validation Split in Meta-Learning?

Yu Bai, Minshuo Chen, Pan Zhou +5

Meta-learning aims to perform fast adaptation on a new task through learning a "prior" from multiple existing tasks. A common practice in meta-learning is to perform a train-valida…

cs.DC2018

Joint Service Caching and Task Offloading for Mobile Edge Computing in Dense Networks

Jie Xu, Lixing Chen, Pan Zhou

Mobile Edge Computing (MEC) pushes computing functionalities away from the centralized cloud to the network edge, thereby meeting the latency requirements of many emerging mobile a…

cs.CE2025

Revisiting the Canonicalization for Fast and Accurate Crystal Tensor Property Prediction

Haowei Hua, Jingwen Yang, Wanyu Lin +1

Predicting the tensor properties of crystalline materials is a fundamental task in materials science. Unlike scalar property prediction, which requires invariance, tensor property…

cs.CV2024

MetaCloak: Preventing Unauthorized Subject-driven Text-to-image Diffusion-based Synthesis via Meta-learning

Yixin Liu, Chenrui Fan, Yutong Dai +3

Text-to-image diffusion models allow seamless generation of personalized images from scant reference photos. Yet, these tools, in the wrong hands, can fabricate misleading or harmf…

cs.DC2021

Gradient Scheduling with Global Momentum for Non-IID Data Distributed Asynchronous Training

Chengjie Li, Ruixuan Li, Haozhao Wang +4

Distributed asynchronous offline training has received widespread attention in recent years because of its high performance on large-scale data and complex models. As data are dist…

cs.LG2026

Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence

Wenzhe Yin, Zehao Xiao, Pan Zhou +4

Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize…

cs.CV2024

MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Shanghua Gao, Pan Zhou, Ming-Ming Cheng +1

Despite its success in image synthesis, we observe that diffusion probabilistic models (DPMs) often lack contextual reasoning ability to learn the relations among object parts in a…

cs.CR2025

Learning from Few Samples: A Novel Approach for High-Quality Malcode Generation

Haijian Ma, Daizong Liu, Xiaowen Cai +2

Intrusion Detection Systems (IDS) play a crucial role in network security defense. However, a significant challenge for IDS in training detection models is the shortage of adequate…

cs.CR2025

BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization

Xueyang Zhou, Guiyao Tie, Guowen Zhang +3

Vision-Language-Action (VLA) models have advanced robotic control by enabling end-to-end decision-making directly from multimodal inputs. However, their tightly coupled architectur…

cs.CV2025

GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding

Dongping Chen, Yue Huang, Siyuan Wu +17

Recently, Multimodal Large Language Models (MLLMs) have been used as agents to control keyboard and mouse inputs by directly perceiving the Graphical User Interface (GUI) and gener…

cs.LG2026

Learning ECG Image Representations via Dual Physiological-Aware Alignments

Hung Manh Pham, Jialu Tang, Aaqib Saeed +3

Electrocardiograms (ECGs) are among the most widely used diagnostic tools for cardiovascular diseases, and a large amount of ECG data worldwide appears only in image form. However,…

cs.LG2024

Provably Efficient Action-Manipulation Attack Against Continuous Reinforcement Learning

Zhi Luo, Xiyuan Yang, Pan Zhou +1

Manipulating the interaction trajectories between the intelligent agent and the environment can control the agent's training and behavior, exposing the potential vulnerabilities of…

cs.CV2022

Backdoor Attacks on Crowd Counting

Yuhua Sun, Tailai Zhang, Xingjun Ma +6

Crowd counting is a regression task that estimates the number of people in a scene image, which plays a vital role in a range of safety-critical applications, such as video surveil…

cond-mat.mtrl-sci2013

Magnetic Properties of Single Transition-Metal Atom Absorbed Graphdiyne and Graphyne Sheet

Junjie He, ShuangYing Ma, Pan Zhou +3

The electronic and magnetic properties of single 3d transition-metal(TM) atom (V, Cr, Mn, Fe, Co, and Ni) adsorbed graphdiyne (GDY) and graphyne (GY) are systematically studied usi…

cs.CV2023

FAT: Feature-Focusing Adversarial Training via Disentanglement of Natural and Perturbed Patterns

Yaguan Qian, Chenyu Zhao, Zhaoquan Gu +5

Deep neural networks (DNNs) are vulnerable to adversarial examples crafted by well-designed perturbations. This could lead to disastrous results on critical applications such as se…

cs.MM2026

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

Yuwen Wang, Tian-Hao Zhang, Minghao Cai +7

Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rath…

cs.CL2024

Cooperative SQL Generation for Segmented Databases By Using Multi-functional LLM Agents

Zhiguang Wu, Fengbin Zhu, Xuequn Shang +2

Text-to-SQL task aims to automatically yield SQL queries according to user text questions. To address this problem, we propose a Cooperative SQL Generation framework based on Multi…

cs.CV2024

Unsupervised Modality Adaptation with Text-to-Image Diffusion Models for Semantic Segmentation

Ruihao Xia, Yu Liang, Peng-Tao Jiang +4

Despite their success, unsupervised domain adaptation methods for semantic segmentation primarily focus on adaptation between image domains and do not utilize other abundant visual…

cs.CV2026

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

Xueyang Zhou, Yangming Xu, Guiyao Tie +5

LIBERO has emerged as a widely adopted benchmark for evaluating Vision-Language-Action (VLA) models; however, its current training and evaluation settings are problematic, often le…

cs.CV2026

Hard-Label Black-Box Attacks on 3D Point Clouds

Daizong Liu, Yunbo Tao, Junhao Dong +4

With the maturity of depth sensors in various 3D safety-critical applications, 3D point cloud models have been shown to be vulnerable to adversarial attacks. Almost all existing 3D…

cs.CL2021

Exploiting Global Contextual Information for Document-level Named Entity Recognition

Zanbo Wang, Wei Wei, Xianling Mao +4

Most existing named entity recognition (NER) approaches are based on sequence labeling models, which focus on capturing the local context dependencies. However, the way of taking o…

cs.AI2025

MMLU-Reason: Benchmarking Multi-Task Multi-modal Language Understanding and Reasoning

Guiyao Tie, Xueyang Zhou, Tianhe Gu +7

Recent advances in Multi-Modal Large Language Models (MLLMs) have enabled unified processing of language, vision, and structured inputs, opening the door to complex tasks such as l…

cs.AI2026

Latent Thought Flow: Efficient Latent Reasoning in Large Language Models

Xiandong Zou, Jing Huang, Jianshu Li +1

Large Language Models (LLMs) increasingly rely on intermediate reasoning, yet explicit Chain-of-Thought (CoT) suffers from a linguistic space bottleneck: each thought must be decod…

cs.CL2021

Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition

Guolin Zheng, Yubei Xiao, Ke Gong +3

Unifying acoustic and linguistic representation learning has become increasingly crucial to transfer the knowledge learned on the abundance of high-resource language data for low-r…

cs.CL2019

An Online Attention-based Model for Speech Recognition

Ruchao Fan, Pan Zhou, Wei Chen +2

Attention-based end-to-end models such as Listen, Attend and Spell (LAS), simplify the whole pipeline of traditional automatic speech recognition (ASR) systems and become popular i…

cs.CV2023

Position-guided Text Prompt for Vision-Language Pre-training

Alex Jinpeng Wang, Pan Zhou, Mike Zheng Shou +1

Vision-Language Pre-Training (VLP) has shown promising capabilities to align image and text pairs, facilitating a broad variety of cross-modal learning tasks. However, we observe t…

cs.LG2026

V3H: View Variation and View Heredity for Incomplete Multi-view Clustering

Xiang Fang, Yuchong Hu, Pan Zhou +1

Real data often appear in the form of multiple incomplete views. Incomplete multi-view clustering is an effective method to integrate these incomplete views. Previous methods only…

cs.CV2025

Towards Understanding Why Data Augmentation Improves Generalization

Jingyang Li, Jiachun Pan, Kim-Chuan Toh +1

Data augmentation is a cornerstone technique in deep learning, widely used to improve model generalization. Traditional methods like random cropping and color jittering, as well as…

cond-mat.mes-hall2026

Layer Edelstein Effect

Binchang Zhou, Pan Zhou, Baoru Pan +3

Electrical control of magnetism represents a fundamental route toward next-generation spintronic functionalities. In this Letter, we introduce a universal current-induced spin phen…

cs.CL2020

Graph-Evolving Meta-Learning for Low-Resource Medical Dialogue Generation

Shuai Lin, Pan Zhou, Xiaodan Liang +4

Human doctors with well-structured medical knowledge can diagnose a disease merely via a few conversations with patients about symptoms. In contrast, existing knowledge-grounded di…

math.OC2018

Faster First-Order Methods for Stochastic Non-Convex Optimization on Riemannian Manifolds

Pan Zhou, Xiao-Tong Yuan, Jiashi Feng

SPIDER (Stochastic Path Integrated Differential EstimatoR) is an efficient gradient estimation technique developed for non-convex stochastic optimization. Although having been show…

cs.LG2023

Towards Understanding Why Mask-Reconstruction Pretraining Helps in Downstream Tasks

Jiachun Pan, Pan Zhou, Shuicheng Yan

For unsupervised pretraining, mask-reconstruction pretraining (MRP) approaches, e.g. MAE and data2vec, randomly mask input patches and then reconstruct the pixels or semantic featu…

cs.AI2025

A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models

Zhongzhan Huang, Shanshan Zhong, Pan Zhou +3

Recently, numerous benchmarks have been developed to evaluate the logical reasoning abilities of large language models (LLMs). However, assessing the equally important creative cap…

cs.CL2024

The Impact of Large Language Models in Academia: from Writing to Speaking

Mingmeng Geng, Caixi Chen, Yanru Wu +3

Large language models (LLMs) are increasingly impacting human society, particularly in textual information. Based on more than 30,000 papers and 1,000 presentations from machine le…

cs.CL2025

CrowdSelect: Synthetic Instruction Data Selection with Multi-LLM Wisdom

Yisen Li, Lingfeng Yang, Wenxuan Shen +4

Distilling advanced Large Language Models' instruction-following capabilities into smaller models using a selected subset has become a mainstream approach in model training. While…

cs.CL2024

Two are better than one: Context window extension with multi-grained self-injection

Wei Han, Pan Zhou, Soujanya Poria +1

The limited context window of contemporary large language models (LLMs) remains a huge barrier to their broader application across various domains. While continual pre-training on…

cs.CL2026

Benchmarking Gaslighting Attacks Against Speech Large Language Models

Jinyang Wu, Bin Zhu, Xiandong Zou +3

As Speech Large Language Models (Speech LLMs) become increasingly integrated into voice-based applications, ensuring their robustness against manipulative or adversarial input beco…

cs.AI2024

Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation

Shanshan Zhong, Zhongzhan Huang, Shanghua Gao +4

Chain-of-Thought (CoT) guides large language models (LLMs) to reason step-by-step, and can motivate their logical reasoning ability. While effective for logical tasks, CoT is not c…

cs.CV2022

Memory-Guided Semantic Learning Network for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Xing Di +3

Temporal sentence grounding (TSG) is crucial and fundamental for video understanding. Although the existing methods train well-designed deep networks with a large amount of data, w…

cs.SE2024

MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Yue Huang, Jiawen Shi, Yuan Li +8

Large language models (LLMs) have garnered significant attention due to their impressive natural language processing (NLP) capabilities. Recently, many studies have focused on the…

cs.CV2023

Transform-Equivariant Consistency Learning for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Jianfeng Dong +6

This paper addresses the temporal sentence grounding (TSG). Although existing methods have made decent achievements in this task, they not only severely rely on abundant video-quer…

cs.CV2021

Towards Adversarial Patch Analysis and Certified Defense against Crowd Counting

Qiming Wu, Zhikang Zou, Pan Zhou +3

Crowd counting has drawn much attention due to its importance in safety-critical surveillance systems. Especially, deep neural network (DNN) methods have significantly reduced esti…

cs.LG2024

Graph Agent Network: Empowering Nodes with Inference Capabilities for Adversarial Resilience

Ao Liu, Wenshan Li, Tao Li +5

End-to-end training with global optimization have popularized graph neural networks (GNNs) for node classification, yet inadvertently introduced vulnerabilities to adversarial edge…

cs.CV2021

Progressively Guide to Attend: An Iterative Alignment Framework for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Pan Zhou

A key solution to temporal sentence grounding (TSG) exists in how to learn effective alignment between vision and language features extracted from an untrimmed video and a sentence…

cond-mat.mtrl-sci2024

General Stacking Theory for Altermagnetism in Bilayer Systems

Baoru Pan, Pan Zhou, Pengbo Lyu +3

Two-dimensional (2D) altermagnetism was recently proposed to be attainable in twisted antiferromagnetic bilayers providing an experimentally feasible approach to realize it in 2D m…

cs.RO2026

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development

Xueyang Zhou, Yihan Sun, Xijie Gong +4

Embodied AI research is increasingly moving beyond single-task, single-environment policy learning toward multi-task, multi-scene, and multi-model settings. This shift substantiall…

cs.LG2023

Iterative Graph Self-Distillation

Hanlin Zhang, Shuai Lin, Weiyang Liu +4

Recently, there has been increasing interest in the challenge of how to discriminatively vectorize graphs. To address this, we propose a method called Iterative Graph Self-Distilla…

cs.AI2026

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

Guiyao Tie, Jiawen Shi, Dingjie Song +20

Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding, hypothesis generation, exper…

cs.CV2017

Jointly Attentive Spatial-Temporal Pooling Networks for Video-based Person Re-Identification

Shuangjie Xu, Yu Cheng, Kang Gu +3

Person Re-Identification (person re-id) is a crucial task as its applications in visual surveillance and human-computer interaction. In this work, we present a novel joint Spatial…

cs.CV2024

Gamba: Marry Gaussian Splatting with Mamba for single view 3D reconstruction

Qiuhong Shen, Zike Wu, Xuanyu Yi +4

We tackle the challenge of efficiently reconstructing a 3D asset from a single image at millisecond speed. Existing methods for single-image 3D reconstruction are primarily based o…

cs.CV2026

Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding

Xiang Fang, Daizong Liu, Wanlong Fang +4

This paper addresses the task of temporal sentence grounding (TSG). Although many respectable works have made decent achievements in this important topic, they severely rely on mas…

cs.NI2021

Cost-efficient and Skew-aware Data Scheduling for Incremental Learning in 5G Network

Lingjun Pu, Xinjing Yuan, Xiaohang Xu +3

To facilitate the emerging applications in 5G networks, mobile network operators will provide many network functions in terms of control and prediction. Recently, they have recogni…

cs.CV2025

LOVA3: Learning to Visual Question Answering, Asking and Assessment

Henry Hengyuan Zhao, Pan Zhou, Difei Gao +2

Question answering, asking, and assessment are three innate human traits crucial for understanding the world and acquiring knowledge. By enhancing these capabilities, humans can mo…

cs.CV2025

Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment

Dongping Chen, Ruoxi Chen, Shu Pu +8

Many real-world user queries (e.g. "How do to make egg fried rice?") could benefit from systems capable of generating responses with both textual steps with accompanying images, si…

cs.CL2024

Can Large Language Models Automatically Jailbreak GPT-4V?

Yuanwei Wu, Yue Huang, Yixin Liu +3

GPT-4V has attracted considerable attention due to its extraordinary capacity for integrating and processing multimodal information. At the same time, its ability of face recogniti…

cs.LG2025

Conda: Column-Normalized Adam for Training Large Language Models Faster

Junjie Wang, Pan Zhou, Yiming Dong +6

Large language models (LLMs) have demonstrated impressive generalization and emergent capabilities, yet their pre-training remains computationally expensive and sensitive to optimi…

cs.NI2016

Near Optimal Adaptive Shortest Path Routing with Stochastic Links States under Adversarial Attack

Pan Zhou, Lin Cheng, Dapeng Oliver Wu

We consider the shortest path routing (SPR) of a network with stochastically time varying link metrics under potential adversarial attacks. Due to potential denial of service attac…

cs.CR2025

Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models

Zenghui Yuan, Yangming Xu, Jiawen Shi +2

Model merging for Large Language Models (LLMs) directly fuses the parameters of different models finetuned on various tasks, creating a unified model for multi-domain tasks. Howeve…

cs.CL2024

Self-Cognition in Large Language Models: An Exploratory Study

Dongping Chen, Jiawen Shi, Yao Wan +3

While Large Language Models (LLMs) have achieved remarkable success across various applications, they also raise concerns regarding self-cognition. In this paper, we perform a pion…

cs.CY2016

On Diffusion-restricted Social Network: A Measurement Study of WeChat Moments

Zhuqi Li, Lin Chen, Yichong Bai +2

WeChat is a mobile messaging application that has 549 million active users as of Q1 2015, and "WeChat Moments" (WM) serves its social-networking function that allows users to post/…