Publications (274)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
Yinuo Liu, Zenghui Yuan, Guiyao Tie +4
Multimodal retrieval-augmented generation (RAG) enhances the visual reasoning capability of vision-language models (VLMs) by dynamically accessing information from external knowled…
Exact Recovery of Tensor Robust Principal Component Analysis under Linear Transforms
Canyi Lu, Pan Zhou
This work studies the Tensor Robust Principal Component Analysis (TRPCA) problem, which aims to exactly recover the low-rank and sparse components from their sum. Our model is moti…
DeepSpace: An Online Deep Learning Framework for Mobile Big Data to Understand Human Mobility Patterns
Xi Ouyang, Chaoyun Zhang, Pan Zhou +2
In the recent years, the rapid spread of mobile device has create the vast amount of mobile data. However, some shallow-structure models such as support vector machine (SVM) have d…
Prompt Injection Attack to Tool Selection in LLM Agents
Jiawen Shi, Zenghui Yuan, Guiyao Tie +3
Tool selection is a key component of LLM agents. A popular approach follows a two-step process - \emph{retrieval} and \emph{selection} - to pick the most appropriate tool from a to…
F2Net: Learning to Focus on the Foreground for Unsupervised Video Object Segmentation
Daizong Liu, Dongdong Yu, Changhu Wang +1
Although deep learning based methods have achieved great progress in unsupervised video object segmentation, difficult scenarios (e.g., visual similarity, occlusions, and appearanc…
Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator
Henry Hengyuan Zhao, Pan Zhou, Mike Zheng Shou
Multimodal Large Language Models (MLLMs) demonstrate exceptional problem-solving capabilities, but few research studies aim to gauge the ability to generate visual instruction tuni…
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
Xingyu Xie, Pan Zhou, Huan Li +2
In deep learning, different kinds of deep networks typically need different optimizers, which have to be chosen after multiple trials, making the training process inefficient. To r…
Towards Scalable and Consistent 3D Editing
Ruihao Xia, Yang Tang, Pan Zhou
3D editing - the task of locally modifying the geometry or appearance of a 3D asset - has wide applications in immersive content creation, digital entertainment, and AR/VR. However…
Theory-Inspired Path-Regularized Differential Network Architecture Search
Pan Zhou, Caiming Xiong, Richard Socher +1
Despite its high search efficiency, differential architecture search (DARTS) often selects network architectures with dominated skip connections which lead to performance degradati…
Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment Localization
Daizong Liu, Xiaoye Qu, Xiao-Yang Liu +3
Query-based moment localization is a new task that localizes the best matched segment in an untrimmed video according to a given sentence query. In this localization task, one shou…
Way to Specialist: Closing Loop Between Specialized LLM and Evolving Domain Knowledge Graph
Yutong Zhang, Lixing Chen, Shenghong Li +6
Large language models (LLMs) have demonstrated exceptional performance across a wide variety of domains. Nonetheless, generalist LLMs continue to fall short in reasoning tasks nece…
Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection
Yuqi Zhou, Lin Lu, Hanchi Sun +2
Jailbreak attacks on large language models (LLMs) involve inducing these models to generate harmful content that violates ethics or laws, posing a significant threat to LLM securit…
Progressive Localization Networks for Language-based Moment Localization
Qi Zheng, Jianfeng Dong, Xiaoye Qu +5
This paper targets the task of language-based video moment localization. The language-based setting of this task allows for an open set of target activities, resulting in a large v…
Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents
Zhejian Yang, Yongchao Chen, Xueyang Zhou +8
Long-horizon robotic manipulation poses significant challenges for autonomous systems, requiring extended reasoning, precise execution, and robust error recovery across complex seq…
MetaFormer Is Actually What You Need for Vision
Weihao Yu, Mi Luo, Pan Zhou +5
Transformers have shown great potential in computer vision tasks. A common belief is their attention-based token mixer module contributes most to their competence. However, recent…
4-bit Shampoo for Memory-Efficient Network Training
Sike Wang, Pan Zhou, Jia Li +1
Second-order optimizers, maintaining a matrix termed a preconditioner, are superior to first-order optimizers in both theory and practice. The states forming the preconditioner and…
Instant3D: Instant Text-to-3D Generation
Ming Li, Pan Zhou, Jia-Wei Liu +4
Text-to-3D generation has attracted much attention from the computer vision community. Existing methods mainly optimize a neural field from scratch for each text prompt, relying on…
ReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment
Wanjiang Weng, Xiaofeng Tan, Hongsong Wang +1
Bilingual text-to-motion generation, which synthesizes 3D human motions from bilingual text inputs, holds immense potential for cross-linguistic applications in gaming, film, and r…
GraphCloak: Safeguarding Task-specific Knowledge within Graph-structured Data from Unauthorized Exploitation
Yixin Liu, Chenrui Fan, Xun Chen +2
As Graph Neural Networks (GNNs) become increasingly prevalent in a variety of fields, from social network analysis to protein-protein interaction studies, growing concerns have eme…
Self-Promoted Supervision for Few-Shot Transformer
Bowen Dong, Pan Zhou, Shuicheng Yan +1
The few-shot learning ability of vision transformers (ViTs) is rarely investigated though heavily desired. In this work, we empirically find that with the same few-shot learning fr…
Automatic channel selection and spatial feature integration for multi-channel speech recognition across various array topologies
Bingshen Mu, Pengcheng Guo, Dake Guo +3
Automatic Speech Recognition (ASR) has shown remarkable progress, yet it still faces challenges in real-world distant scenarios across various array topologies each with multiple r…
BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT
Jiawen Shi, Yixin Liu, Pan Zhou +1
Recently, ChatGPT has gained significant attention in research due to its ability to interact with humans effectively. The core idea behind this model is reinforcement learning (RL…
Communication-efficient Decentralized Machine Learning over Heterogeneous Networks
Pan Zhou, Qian Lin, Dumitrel Loghin +3
In the last few years, distributed machine learning has been usually executed over heterogeneous networks such as a local area network within a multi-tenant cluster or a wide area…
U2-KWS: Unified Two-pass Open-vocabulary Keyword Spotting with Keyword Bias
Ao Zhang, Pan Zhou, Kaixun Huang +3
Open-vocabulary keyword spotting (KWS), which allows users to customize keywords, has attracted increasingly more interest. However, existing methods based on acoustic models and p…
Differentially Private Distributed Online Learning
Chencheng Li, Pan Zhou
Online learning has been in the spotlight from the machine learning society for a long time. To handle massive data in Big Data era, one single learner could never efficiently fini…
What Makes Good Collaborative Views? Contrastive Mutual Information Maximization for Multi-Agent Perception
Wanfang Su, Lixing Chen, Yang Bai +4
Multi-agent perception (MAP) allows autonomous systems to understand complex environments by interpreting data from multiple sources. This paper investigates intermediate collabora…
Exploring RNN-Transducer for Chinese Speech Recognition
Senmao Wang, Pan Zhou, Wei Chen +2
End-to-end approaches have drawn much attention recently for significantly simplifying the construction of an automatic speech recognition (ASR) system. RNN transducer (RNN-T) is o…
High-Throughput Discovery of Two-Dimensional Materials Exhibiting Strong Rashba-Edelstein effect
Binchang Zhou, Baoru Pan, Pan Zhou +3
The Rashba-Edelstein effect (REE), which generates spin accumulation under an applied electric current, quantifies charge-to-spin conversion (CSC) efficiency in non-centrosymmetric…
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
Shijing Hu, Jingyang Li, Xingyu Xie +3
Speculative decoding accelerates inference in large language models (LLMs) by generating multiple draft tokens simultaneously. However, existing methods often struggle with token m…
A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
Daizong Liu, Mingyu Yang, Xiaoye Qu +3
With the significant development of large models in recent years, Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal u…
Transfer Heterogeneous Knowledge Among Peer-to-Peer Teammates: A Model Distillation Approach
Zeyue Xue, Shuang Luo, Chao Wu +3
Peer-to-peer knowledge transfer in distributed environments has emerged as a promising method since it could accelerate learning and improve team-wide performance without relying o…
The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023
He Wang, Pengcheng Guo, Wei Chen +2
This paper delineates the visual speech recognition (VSR) system introduced by the NPU-ASLP-LiAuto (Team 237) in the first Chinese Continuous Visual Speech Recognition Challenge (C…
IBNorm: Information-Bottleneck Inspired Normalization for Representation Learning
Xiandong Zou, Jia Li, Xiaotong Yuan +1
Normalization is fundamental to deep learning, but existing approaches such as BatchNorm, LayerNorm, and RMSNorm are variance-centric by enforcing zero mean and unit variance, stab…
Temporal-Guided Visual Foundation Models for Event-Based Vision
Ruihao Xia, Junhong Cai, Luziwei Leng +5
Event cameras offer unique advantages for vision tasks in challenging environments, yet processing asynchronous event streams remains an open challenge. While existing methods rely…
LiAuto-GeoX: Efficient Grounded Driving Transformer
Jiawei Lian, Haoyi Sun, Yang Wu +8
Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for autonomous driving remains an ope…
Friendly Sharpness-Aware Minimization
Tao Li, Pan Zhou, Zhengbao He +2
Sharpness-Aware Minimization (SAM) has been instrumental in improving deep neural network training by minimizing both training loss and loss sharpness. Despite the practical succes…
STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition
Ming Li, Xiangyu Xu, Hehe Fan +7
Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks…
A Theory-Driven Self-Labeling Refinement Method for Contrastive Representation Learning
Pan Zhou, Caiming Xiong, Xiao-Tong Yuan +1
For an image query, unsupervised contrastive learning labels crops of the same image as positives, and other image crops as negatives. Although intuitive, such a native label assig…
3DHacker: Spectrum-based Decision Boundary Generation for Hard-label 3D Point Cloud Attack
Yunbo Tao, Daizong Liu, Pan Zhou +3
With the maturity of depth sensors, the vulnerability of 3D point cloud models has received increasing attention in various applications such as autonomous driving and robot naviga…
Backdoor Attacks to Pre-trained Unified Foundation Models
Zenghui Yuan, Yixin Liu, Kai Zhang +2
The rise of pre-trained unified foundation models breaks down the barriers between different modalities and tasks, providing comprehensive support to users with unified architectur…
How Important is the Train-Validation Split in Meta-Learning?
Yu Bai, Minshuo Chen, Pan Zhou +5
Meta-learning aims to perform fast adaptation on a new task through learning a "prior" from multiple existing tasks. A common practice in meta-learning is to perform a train-valida…
Joint Service Caching and Task Offloading for Mobile Edge Computing in Dense Networks
Jie Xu, Lixing Chen, Pan Zhou
Mobile Edge Computing (MEC) pushes computing functionalities away from the centralized cloud to the network edge, thereby meeting the latency requirements of many emerging mobile a…
Revisiting the Canonicalization for Fast and Accurate Crystal Tensor Property Prediction
Haowei Hua, Jingwen Yang, Wanyu Lin +1
Predicting the tensor properties of crystalline materials is a fundamental task in materials science. Unlike scalar property prediction, which requires invariance, tensor property…
MetaCloak: Preventing Unauthorized Subject-driven Text-to-image Diffusion-based Synthesis via Meta-learning
Yixin Liu, Chenrui Fan, Yutong Dai +3
Text-to-image diffusion models allow seamless generation of personalized images from scant reference photos. Yet, these tools, in the wrong hands, can fabricate misleading or harmf…
Gradient Scheduling with Global Momentum for Non-IID Data Distributed Asynchronous Training
Chengjie Li, Ruixuan Li, Haozhao Wang +4
Distributed asynchronous offline training has received widespread attention in recent years because of its high performance on large-scale data and complex models. As data are dist…
Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence
Wenzhe Yin, Zehao Xiao, Pan Zhou +4
Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize…
MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer
Shanghua Gao, Pan Zhou, Ming-Ming Cheng +1
Despite its success in image synthesis, we observe that diffusion probabilistic models (DPMs) often lack contextual reasoning ability to learn the relations among object parts in a…
Learning from Few Samples: A Novel Approach for High-Quality Malcode Generation
Haijian Ma, Daizong Liu, Xiaowen Cai +2
Intrusion Detection Systems (IDS) play a crucial role in network security defense. However, a significant challenge for IDS in training detection models is the shortage of adequate…
BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization
Xueyang Zhou, Guiyao Tie, Guowen Zhang +3
Vision-Language-Action (VLA) models have advanced robotic control by enabling end-to-end decision-making directly from multimodal inputs. However, their tightly coupled architectur…
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
Dongping Chen, Yue Huang, Siyuan Wu +17
Recently, Multimodal Large Language Models (MLLMs) have been used as agents to control keyboard and mouse inputs by directly perceiving the Graphical User Interface (GUI) and gener…
Learning ECG Image Representations via Dual Physiological-Aware Alignments
Hung Manh Pham, Jialu Tang, Aaqib Saeed +3
Electrocardiograms (ECGs) are among the most widely used diagnostic tools for cardiovascular diseases, and a large amount of ECG data worldwide appears only in image form. However,…
Provably Efficient Action-Manipulation Attack Against Continuous Reinforcement Learning
Zhi Luo, Xiyuan Yang, Pan Zhou +1
Manipulating the interaction trajectories between the intelligent agent and the environment can control the agent's training and behavior, exposing the potential vulnerabilities of…
Backdoor Attacks on Crowd Counting
Yuhua Sun, Tailai Zhang, Xingjun Ma +6
Crowd counting is a regression task that estimates the number of people in a scene image, which plays a vital role in a range of safety-critical applications, such as video surveil…
Magnetic Properties of Single Transition-Metal Atom Absorbed Graphdiyne and Graphyne Sheet
Junjie He, ShuangYing Ma, Pan Zhou +3
The electronic and magnetic properties of single 3d transition-metal(TM) atom (V, Cr, Mn, Fe, Co, and Ni) adsorbed graphdiyne (GDY) and graphyne (GY) are systematically studied usi…
FAT: Feature-Focusing Adversarial Training via Disentanglement of Natural and Perturbed Patterns
Yaguan Qian, Chenyu Zhao, Zhaoquan Gu +5
Deep neural networks (DNNs) are vulnerable to adversarial examples crafted by well-designed perturbations. This could lead to disastrous results on critical applications such as se…
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models
Yuwen Wang, Tian-Hao Zhang, Minghao Cai +7
Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rath…
Cooperative SQL Generation for Segmented Databases By Using Multi-functional LLM Agents
Zhiguang Wu, Fengbin Zhu, Xuequn Shang +2
Text-to-SQL task aims to automatically yield SQL queries according to user text questions. To address this problem, we propose a Cooperative SQL Generation framework based on Multi…
Unsupervised Modality Adaptation with Text-to-Image Diffusion Models for Semantic Segmentation
Ruihao Xia, Yu Liang, Peng-Tao Jiang +4
Despite their success, unsupervised domain adaptation methods for semantic segmentation primarily focus on adaptation between image domains and do not utilize other abundant visual…
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
Xueyang Zhou, Yangming Xu, Guiyao Tie +5
LIBERO has emerged as a widely adopted benchmark for evaluating Vision-Language-Action (VLA) models; however, its current training and evaluation settings are problematic, often le…
Hard-Label Black-Box Attacks on 3D Point Clouds
Daizong Liu, Yunbo Tao, Junhao Dong +4
With the maturity of depth sensors in various 3D safety-critical applications, 3D point cloud models have been shown to be vulnerable to adversarial attacks. Almost all existing 3D…
Exploiting Global Contextual Information for Document-level Named Entity Recognition
Zanbo Wang, Wei Wei, Xianling Mao +4
Most existing named entity recognition (NER) approaches are based on sequence labeling models, which focus on capturing the local context dependencies. However, the way of taking o…
MMLU-Reason: Benchmarking Multi-Task Multi-modal Language Understanding and Reasoning
Guiyao Tie, Xueyang Zhou, Tianhe Gu +7
Recent advances in Multi-Modal Large Language Models (MLLMs) have enabled unified processing of language, vision, and structured inputs, opening the door to complex tasks such as l…
Latent Thought Flow: Efficient Latent Reasoning in Large Language Models
Xiandong Zou, Jing Huang, Jianshu Li +1
Large Language Models (LLMs) increasingly rely on intermediate reasoning, yet explicit Chain-of-Thought (CoT) suffers from a linguistic space bottleneck: each thought must be decod…
Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition
Guolin Zheng, Yubei Xiao, Ke Gong +3
Unifying acoustic and linguistic representation learning has become increasingly crucial to transfer the knowledge learned on the abundance of high-resource language data for low-r…
An Online Attention-based Model for Speech Recognition
Ruchao Fan, Pan Zhou, Wei Chen +2
Attention-based end-to-end models such as Listen, Attend and Spell (LAS), simplify the whole pipeline of traditional automatic speech recognition (ASR) systems and become popular i…
Position-guided Text Prompt for Vision-Language Pre-training
Alex Jinpeng Wang, Pan Zhou, Mike Zheng Shou +1
Vision-Language Pre-Training (VLP) has shown promising capabilities to align image and text pairs, facilitating a broad variety of cross-modal learning tasks. However, we observe t…
V3H: View Variation and View Heredity for Incomplete Multi-view Clustering
Xiang Fang, Yuchong Hu, Pan Zhou +1
Real data often appear in the form of multiple incomplete views. Incomplete multi-view clustering is an effective method to integrate these incomplete views. Previous methods only…
Towards Understanding Why Data Augmentation Improves Generalization
Jingyang Li, Jiachun Pan, Kim-Chuan Toh +1
Data augmentation is a cornerstone technique in deep learning, widely used to improve model generalization. Traditional methods like random cropping and color jittering, as well as…
Layer Edelstein Effect
Binchang Zhou, Pan Zhou, Baoru Pan +3
Electrical control of magnetism represents a fundamental route toward next-generation spintronic functionalities. In this Letter, we introduce a universal current-induced spin phen…
Graph-Evolving Meta-Learning for Low-Resource Medical Dialogue Generation
Shuai Lin, Pan Zhou, Xiaodan Liang +4
Human doctors with well-structured medical knowledge can diagnose a disease merely via a few conversations with patients about symptoms. In contrast, existing knowledge-grounded di…
Faster First-Order Methods for Stochastic Non-Convex Optimization on Riemannian Manifolds
Pan Zhou, Xiao-Tong Yuan, Jiashi Feng
SPIDER (Stochastic Path Integrated Differential EstimatoR) is an efficient gradient estimation technique developed for non-convex stochastic optimization. Although having been show…
Towards Understanding Why Mask-Reconstruction Pretraining Helps in Downstream Tasks
Jiachun Pan, Pan Zhou, Shuicheng Yan
For unsupervised pretraining, mask-reconstruction pretraining (MRP) approaches, e.g. MAE and data2vec, randomly mask input patches and then reconstruct the pixels or semantic featu…
A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models
Zhongzhan Huang, Shanshan Zhong, Pan Zhou +3
Recently, numerous benchmarks have been developed to evaluate the logical reasoning abilities of large language models (LLMs). However, assessing the equally important creative cap…
The Impact of Large Language Models in Academia: from Writing to Speaking
Mingmeng Geng, Caixi Chen, Yanru Wu +3
Large language models (LLMs) are increasingly impacting human society, particularly in textual information. Based on more than 30,000 papers and 1,000 presentations from machine le…
CrowdSelect: Synthetic Instruction Data Selection with Multi-LLM Wisdom
Yisen Li, Lingfeng Yang, Wenxuan Shen +4
Distilling advanced Large Language Models' instruction-following capabilities into smaller models using a selected subset has become a mainstream approach in model training. While…
Two are better than one: Context window extension with multi-grained self-injection
Wei Han, Pan Zhou, Soujanya Poria +1
The limited context window of contemporary large language models (LLMs) remains a huge barrier to their broader application across various domains. While continual pre-training on…
Benchmarking Gaslighting Attacks Against Speech Large Language Models
Jinyang Wu, Bin Zhu, Xiandong Zou +3
As Speech Large Language Models (Speech LLMs) become increasingly integrated into voice-based applications, ensuring their robustness against manipulative or adversarial input beco…
Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
Shanshan Zhong, Zhongzhan Huang, Shanghua Gao +4
Chain-of-Thought (CoT) guides large language models (LLMs) to reason step-by-step, and can motivate their logical reasoning ability. While effective for logical tasks, CoT is not c…
Memory-Guided Semantic Learning Network for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Xing Di +3
Temporal sentence grounding (TSG) is crucial and fundamental for video understanding. Although the existing methods train well-designed deep networks with a large amount of data, w…
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Yue Huang, Jiawen Shi, Yuan Li +8
Large language models (LLMs) have garnered significant attention due to their impressive natural language processing (NLP) capabilities. Recently, many studies have focused on the…
Transform-Equivariant Consistency Learning for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Jianfeng Dong +6
This paper addresses the temporal sentence grounding (TSG). Although existing methods have made decent achievements in this task, they not only severely rely on abundant video-quer…
Towards Adversarial Patch Analysis and Certified Defense against Crowd Counting
Qiming Wu, Zhikang Zou, Pan Zhou +3
Crowd counting has drawn much attention due to its importance in safety-critical surveillance systems. Especially, deep neural network (DNN) methods have significantly reduced esti…
Graph Agent Network: Empowering Nodes with Inference Capabilities for Adversarial Resilience
Ao Liu, Wenshan Li, Tao Li +5
End-to-end training with global optimization have popularized graph neural networks (GNNs) for node classification, yet inadvertently introduced vulnerabilities to adversarial edge…
Progressively Guide to Attend: An Iterative Alignment Framework for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Pan Zhou
A key solution to temporal sentence grounding (TSG) exists in how to learn effective alignment between vision and language features extracted from an untrimmed video and a sentence…
General Stacking Theory for Altermagnetism in Bilayer Systems
Baoru Pan, Pan Zhou, Pengbo Lyu +3
Two-dimensional (2D) altermagnetism was recently proposed to be attainable in twisted antiferromagnetic bilayers providing an experimentally feasible approach to realize it in 2D m…
EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development
Xueyang Zhou, Yihan Sun, Xijie Gong +4
Embodied AI research is increasingly moving beyond single-task, single-environment policy learning toward multi-task, multi-scene, and multi-model settings. This shift substantiall…
Iterative Graph Self-Distillation
Hanlin Zhang, Shuai Lin, Weiyang Liu +4
Recently, there has been increasing interest in the challenge of how to discriminatively vectorize graphs. To address this, we propose a method called Iterative Graph Self-Distilla…
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
Guiyao Tie, Jiawen Shi, Dingjie Song +20
Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding, hypothesis generation, exper…
Jointly Attentive Spatial-Temporal Pooling Networks for Video-based Person Re-Identification
Shuangjie Xu, Yu Cheng, Kang Gu +3
Person Re-Identification (person re-id) is a crucial task as its applications in visual surveillance and human-computer interaction. In this work, we present a novel joint Spatial…
Gamba: Marry Gaussian Splatting with Mamba for single view 3D reconstruction
Qiuhong Shen, Zike Wu, Xuanyu Yi +4
We tackle the challenge of efficiently reconstructing a 3D asset from a single image at millisecond speed. Existing methods for single-image 3D reconstruction are primarily based o…
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding
Xiang Fang, Daizong Liu, Wanlong Fang +4
This paper addresses the task of temporal sentence grounding (TSG). Although many respectable works have made decent achievements in this important topic, they severely rely on mas…
Cost-efficient and Skew-aware Data Scheduling for Incremental Learning in 5G Network
Lingjun Pu, Xinjing Yuan, Xiaohang Xu +3
To facilitate the emerging applications in 5G networks, mobile network operators will provide many network functions in terms of control and prediction. Recently, they have recogni…
LOVA3: Learning to Visual Question Answering, Asking and Assessment
Henry Hengyuan Zhao, Pan Zhou, Difei Gao +2
Question answering, asking, and assessment are three innate human traits crucial for understanding the world and acquiring knowledge. By enhancing these capabilities, humans can mo…
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
Dongping Chen, Ruoxi Chen, Shu Pu +8
Many real-world user queries (e.g. "How do to make egg fried rice?") could benefit from systems capable of generating responses with both textual steps with accompanying images, si…
Can Large Language Models Automatically Jailbreak GPT-4V?
Yuanwei Wu, Yue Huang, Yixin Liu +3
GPT-4V has attracted considerable attention due to its extraordinary capacity for integrating and processing multimodal information. At the same time, its ability of face recogniti…
Conda: Column-Normalized Adam for Training Large Language Models Faster
Junjie Wang, Pan Zhou, Yiming Dong +6
Large language models (LLMs) have demonstrated impressive generalization and emergent capabilities, yet their pre-training remains computationally expensive and sensitive to optimi…
Near Optimal Adaptive Shortest Path Routing with Stochastic Links States under Adversarial Attack
Pan Zhou, Lin Cheng, Dapeng Oliver Wu
We consider the shortest path routing (SPR) of a network with stochastically time varying link metrics under potential adversarial attacks. Due to potential denial of service attac…
Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
Zenghui Yuan, Yangming Xu, Jiawen Shi +2
Model merging for Large Language Models (LLMs) directly fuses the parameters of different models finetuned on various tasks, creating a unified model for multi-domain tasks. Howeve…
Self-Cognition in Large Language Models: An Exploratory Study
Dongping Chen, Jiawen Shi, Yao Wan +3
While Large Language Models (LLMs) have achieved remarkable success across various applications, they also raise concerns regarding self-cognition. In this paper, we perform a pion…
On Diffusion-restricted Social Network: A Measurement Study of WeChat Moments
Zhuqi Li, Lin Chen, Yichong Bai +2
WeChat is a mobile messaging application that has 549 million active users as of Q1 2015, and "WeChat Moments" (WM) serves its social-networking function that allows users to post/…