papers

Publications (98)

cs.CL2026

KaLM: Knowledge-aligned Autoregressive Language Modeling via Dual-view Knowledge Graph Contrastive Learning

Peng Yu, Cheng Deng, Beiya Dai +2

Autoregressive large language models (LLMs) pre-trained by next token prediction are inherently proficient in generative tasks. However, their performance on knowledge-driven tasks…

cs.CV2026

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents

Jiahua Li, Zhanhe Zhang, Chenghao Xu +4

Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Although existing Large Language Model…

cs.CL2025

PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing

Cheng Deng, Luoyang Sun, Jiwen Jiang +10

While scaling laws have been continuously validated in large language models (LLMs) with increasing model parameters, the inherent tension between the inference demands of LLMs and…

cs.RO2026

DSCD-Nav: Dual-Stance Cooperative Debate for Object Navigation

Weitao An, Qi Liu, Chenghao Xu +4

Adaptive navigation in unfamiliar indoor environments is crucial for household service robots. Despite advances in zero-shot perception and reasoning from vision-language models, e…

cs.CV2026

SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs

Siting Wang, Minnan Pei, Luoyang Sun +6

Humans can imagine and manipulate visual images mentally, a capability known as spatial visualization. While many multi-modal benchmarks assess reasoning on visible visual informat…

math.OC2021

Faster Stochastic Quasi-Newton Methods

Qingsong Zhang, Feihu Huang, Cheng Deng +1

Stochastic optimization methods have become a class of popular optimization tools in machine learning. Especially, stochastic gradient descent (SGD) has been widely used for machin…

cs.RO2026

Semantic Evidence Regulation via Relational Bias for Zero-Shot Object Navigation

Weitao An, Chenghao Xu, Xu Yang +1

Object navigation requires an embodied agent to locate a target object in an unknown environment through visual observations. Existing zero-shot methods typically leverage open-voc…

cs.CV2025

AStF: Motion Style Transfer via Adaptive Statistics Fusor

Hanmo Chen, Chenghao Xu, Jiexi Yan +1

Human motion style transfer allows characters to appear less rigidity and more realism with specific style. Traditional arbitrary image style transfer typically process mean and va…

cs.CL2025

LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues

Haoyang Li, Zhanchao Xu, Yiming Li +9

Multi-turn dialogues are essential in many real-world applications of large language models, such as chatbots and virtual assistants. As conversation histories become longer, exist…

cs.CL2023

Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus

Tianhang Zhang, Lin Qiu, Qipeng Guo +6

Large Language Models (LLMs) have gained significant popularity for their impressive performance across diverse fields. However, LLMs are prone to hallucinate untruthful or nonsens…

cs.LG2019

Deep Spectral Clustering using Dual Autoencoder Network

Xu Yang, Cheng Deng, Feng Zheng +2

The clustering methods have recently absorbed even-increasing attention in learning and vision. Deep clustering combines embedding and clustering together to obtain optimal embeddi…

cs.CV2025

A Tale of Two Experts: Cooperative Learning for Source-Free Unsupervised Domain Adaptation

Jiaping Yu, Muli Yang, Jiapeng Ji +2

Source-Free Unsupervised Domain Adaptation (SFUDA) addresses the realistic challenge of adapting a source-trained model to a target domain without access to the source data, driven…

cs.LG2017

Theoretic Analysis and Extremely Easy Algorithms for Domain Adaptive Feature Learning

Wenhao Jiang, Cheng Deng, Wei Liu +3

Domain adaptation problems arise in a variety of applications, where a training dataset from the \textit{source} domain and a test dataset from the \textit{target} domain typically…

cs.CV2025

Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction

Guangtao Lyu, Xinyi Cheng, Chenghao Xu +7

Large Vision-Language Models (LVLMs) have shown remarkable capabilities, yet hallucinations remain a persistent challenge. This work presents a systematic analysis of the internal…

cs.CV2026

Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion

Hanmo Chen, Chenghao Xu, Xu Yang +2

Video generation is pivotal to digital media creation, and recent advances in autoregressive video generation have markedly enhanced the efficiency of real-time video synthesis. Ho…

cs.CV2026

Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering

Guangtao Lyu, Xinyi Cheng, Qi Liu +5

LVLMs achieve remarkable multimodal understanding and generation but remain susceptible to hallucinations. Existing mitigation methods predominantly focus on output-level adjustmen…

cs.CV2020

Fewer is More: A Deep Graph Metric Learning Perspective Using Fewer Proxies

Yuehua Zhu, Muli Yang, Cheng Deng +1

Deep metric learning plays a key role in various machine learning tasks. Most of the previous works have been confined to sampling from a mini-batch, which cannot precisely charact…

cs.LG2021

Secure Bilevel Asynchronous Vertical Federated Learning with Backward Updating

Qingsong Zhang, Bin Gu, Cheng Deng +1

Vertical federated learning (VFL) attracts increasing attention due to the emerging demands of multi-party collaborative modeling and concerns of privacy leakage. In the real VFL a…

cs.LG2022

Desirable Companion for Vertical Federated Learning: New Zeroth-Order Gradient Based Algorithm

Qingsong Zhang, Bin Gu, Zhiyuan Dang +2

Vertical federated learning (VFL) attracts increasing attention due to the emerging demands of multi-party collaborative modeling and concerns of privacy leakage. A complete list o…

cs.CV2026

Not All Pixels Are Equal: Pixel-wise Meta-Learning for Medical Segmentation with Noisy Labels

Chenyu Mu, Guihai Chen, Xun Yang +2

Medical image segmentation is crucial for clinical applications, but it is frequently disrupted by noisy annotations and ambiguous anatomical boundaries, limiting its application i…

cs.LG2026

Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs

Luoyang Sun, Jiwen Jiang, Yifeng Ding +9

Vision-Language-Action Models (VLAs) have emerged as a key paradigm of Physical AI and are increasingly deployed in autonomous vehicles, robots, and smart spaces. In these resource…

cs.CV2024

Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons

Zhe Xu, Kun Wei, Xu Yang +1

Human dance generation (HDG) aims to synthesize realistic videos from images and sequences of driving poses. Despite great success, existing methods are limited to generating video…

cs.CV2019

Semantic Adversarial Network with Multi-scale Pyramid Attention for Video Classification

De Xie, Cheng Deng, Hao Wang +2

Two-stream architecture have shown strong performance in video classification task. The key idea is to learn spatio-temporal features by fusing convolutional networks spatially and…

cs.CV2019

Deep Multi-scale Discriminative Networks for Double JPEG Compression Forensics

Cheng Deng, Zhao Li, Xinbo Gao +1

As JPEG is the most widely used image format, the importance of tampering detection for JPEG images in blind forensics is self-evident. In this area, extracting effective statistic…

cs.CV2019

Fully-Featured Attribute Transfer

De Xie, Muli Yang, Cheng Deng +2

Image attribute transfer aims to change an input image to a target one with expected attributes, which has received significant attention in recent years. However, most of the exis…

cs.CV2023

Environment-Invariant Curriculum Relation Learning for Fine-Grained Scene Graph Generation

Yukuan Min, Aming Wu, Cheng Deng

The scene graph generation (SGG) task is designed to identify the predicates based on the subject-object pairs.However,existing datasets generally include two imbalance cases: one…

cs.CL2023

K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization

Cheng Deng, Tianhang Zhang, Zhongmou He +9

Large language models (LLMs) have achieved great success in general domains of natural language processing. In this paper, we bring LLMs to the realm of geoscience with the objecti…

cs.LG2021

AsySQN: Faster Vertical Federated Learning Algorithms with Better Computation Resource Utilization

Qingsong Zhang, Bin Gu, Cheng Deng +4

Vertical federated learning (VFL) is an effective paradigm of training the emerging cross-organizational (e.g., different corporations, companies and organizations) collaborative l…

cs.CL2025

NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens

Cunxiang Wang, Ruoxi Ning, Boqi Pan +8

Recent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of…

cs.DL2024

AceMap: Knowledge Discovery through Academic Graph

Xinbing Wang, Luoyi Fu, Xiaoying Gan +23

The exponential growth of scientific literature requires effective management and extraction of valuable insights. While existing scientific search engines excel at delivering sear…

cs.CV2019

Shared Predictive Cross-Modal Deep Quantization

Erkun Yang, Cheng Deng, Chao Li +3

With explosive growth of data volume and ever-increasing diversity of data modalities, cross-modal similarity search, which conducts nearest neighbor search across different modali…

cs.LG2021

Group Contrastive Self-Supervised Learning on Graphs

Xinyi Xu, Cheng Deng, Yaochen Xie +1

We study self-supervised learning on graphs using contrastive methods. A general scheme of prior methods is to optimize two-view representations of input graphs. In many studies, a…

cs.LG2024

Label Learning Method Based on Tensor Projection

Jing Li, Quanxue Gao, Qianqian Wang +2

Multi-view clustering method based on anchor graph has been widely concerned due to its high efficiency and effectiveness. In order to avoid post-processing, most of the existing a…

cs.CV2022

MSR: Making Self-supervised learning Robust to Aggressive Augmentations

Yingbin Bai, Erkun Yang, Zhaoqing Wang +5

Most recent self-supervised learning methods learn visual representation by contrasting different augmented views of images. Compared with supervised learning, more aggressive augm…

cs.CL2023

PK-Chat: Pointer Network Guided Knowledge Driven Generative Dialogue Model

Cheng Deng, Bo Tong, Luoyi Fu +4

In the research of end-to-end dialogue systems, using real-world knowledge to generate natural, fluent, and human-like utterances with correct answers is crucial. However, domain-s…

cs.IR2023

Covidia: COVID-19 Interdisciplinary Academic Knowledge Graph

Cheng Deng, Jiaxin Ding, Luoyi Fu +3

The pandemic of COVID-19 has inspired extensive works across different research fields. Existing literature and knowledge platforms on COVID-19 only focus on collecting papers on b…

cs.CV2025

Towards Arbitrary Motion Completing via Hierarchical Continuous Representation

Chenghao Xu, Guangtao Lyu, Qi Liu +3

Physical motions are inherently continuous, and higher camera frame rates typically contribute to improved smoothness and temporal coherence. For the first time, we explore continu…

cs.LG2026

Rotation Control Unlearning: Quantifying and Controlling Continuous Unlearning for LLM with The Cognitive Rotation Space

Xiang Zhang, Kun Wei, Xu Yang +3

As Large Language Models (LLMs) become increasingly prevalent, their security vulnerabilities have already drawn attention. Machine unlearning is introduced to seek to mitigate the…

cs.LG2024

Self-Supervised Graph Embedding Clustering

Fangfang Li, Quanxue Gao, Cheng Deng +1

The K-means one-step dimensionality reduction clustering method has made some progress in addressing the curse of dimensionality in clustering tasks. However, it combines the K-mea…

cs.MA2025

MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model Framework

Qirui Mi, Mengyue Yang, Xiangning Yu +6

Simulating collective decision-making involves more than aggregating individual behaviors; it emerges from dynamic interactions among individuals. While large language models (LLMs…

cs.CL2025

AceParse: A Comprehensive Dataset with Diverse Structured Texts for Academic Literature Parsing

Huawei Ji, Cheng Deng, Bo Xue +6

With the development of data-centric AI, the focus has shifted from model-driven approaches to improving data quality. Academic literature, as one of the crucial types, is predomin…

cs.CV2019

Towards Visual Feature Translation

Jie Hu, Rongrong Ji, Hong Liu +3

Most existing visual search systems are deployed based upon fixed kinds of visual features, which prohibits the feature reusing across different systems or when upgrading systems w…

cs.CV2023

Invisible Backdoor Attack with Dynamic Triggers against Person Re-identification

Wenli Sun, Xinyang Jiang, Shuguang Dou +4

In recent years, person Re-identification (ReID) has rapidly progressed with wide real-world applications, but also poses significant risks of adversarial attacks. In this paper, w…

cs.LG2026

ContextPilot: Fast Long-Context Inference via Context Reuse

Yinsicheng Jiang, Yeqi Huang, Liang Cheng +3

AI applications increasingly depend on long-context inference, where LLMs consume substantial context to support stronger reasoning. Common examples include retrieval-augmented gen…

cs.CV2021

Adaptive Hierarchical Similarity Metric Learning with Noisy Labels

Jiexi Yan, Lei Luo, Cheng Deng +1

Deep Metric Learning (DML) plays a critical role in various machine learning tasks. However, most existing deep metric learning methods with binary similarity are sensitive to nois…

cs.CV2021

Domain-Smoothing Network for Zero-Shot Sketch-Based Image Retrieval

Zhipeng Wang, Hao Wang, Jiexi Yan +2

Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is a novel cross-modal retrieval task, where abstract sketches are used as queries to retrieve natural images under zero-shot scena…

cs.CL2024

Good Idea or Not, Representation of LLM Could Tell

Yi Xu, Bo Xue, Shuqian Sheng +6

In the ever-expanding landscape of academic research, the proliferation of ideas presents a significant challenge for researchers: discerning valuable ideas from the less impactful…

cs.CV2026

Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning

Hanmo Chen, Guangtao Lyu, Chenghao Xu +3

As a foundational task in human-centric cross-modal intelligence, motion-language retrieval aims to bridge the semantic gap between natural language and human motion, enabling intu…

cs.CV2026

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation

Chenyu Mu, Xin He, Qu Yang +13

Recent advances in video generation have produced models capable of synthesizing stunning visual content from simple text prompts. However, these models struggle to generate long-f…

cs.LG2024

DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning

Siyuan Guo, Cheng Deng, Ying Wen +3

In this work, we investigate the potential of large language models (LLMs) based agents to automate data science tasks, with the goal of comprehending task requirements, then build…

cs.CV2020

Incremental Embedding Learning via Zero-Shot Translation

Kun Wei, Cheng Deng, Xu Yang +1

Modern deep learning methods have achieved great success in machine learning and computer vision fields by learning a set of pre-defined datasets. Howerver, these methods perform u…

cs.CV2022

Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning

Xiangyu Li, Xu Yang, Kun Wei +2

Compositional Zero-Shot Learning (CZSL) aims to recognize unseen compositions formed from seen state and object during training. Since the same state may be various in the visual a…

cs.LG2024

Interpretable Multi-View Clustering Based on Anchor Graph Tensor Factorization

Rui Wang, Jing Li, Quanxue Gao +1

The clustering method based on the anchor graph has gained significant attention due to its exceptional clustering performance and ability to process large-scale data. One common a…

cs.CV2026

Hierarchical Dual-Subspace Decoupling for Continual Learning in Vision-Language Models

Mengxin Qin, Xiang Zhang, Kun Wei +2

Class-incremental learning aims to continuously acquire new knowledge while preserving previously learned information, thereby mitigating catastrophic forgetting. Existing methods…

cs.CL2024

GeoGalactica: A Scientific Large Language Model in Geoscience

Zhouhan Lin, Cheng Deng, Le Zhou +18

Large language models (LLMs) have achieved huge success for their general knowledge and ability to solve a wide spectrum of tasks in natural language processing (NLP). Due to their…

cs.LG2023

FMGNN: Fused Manifold Graph Neural Network

Cheng Deng, Fan Xu, Jiaxing Ding +3

Graph representation learning has been widely studied and demonstrated effectiveness in various graph tasks. Most existing works embed graph data in the Euclidean space, while rece…

cs.CV2021

Towards Improved and Interpretable Deep Metric Learning via Attentive Grouping

Xinyi Xu, Zhengyang Wang, Cheng Deng +2

Grouping has been commonly used in deep metric learning for computing diverse features. However, current methods are prone to overfitting and lack interpretability. In this work, w…

cs.AI2026

Compensating Visual Insufficiency with Stratified Language Guidance for Long-Tail Class Incremental Learning

Xi Wang, Xu Yang, Donghao Sun +1

Long-tail class incremental learning (LT CIL) remains highly challenging because the scarcity of samples in tail classes not only hampers their learning but also exacerbates catast…

cs.LG2026

Revitalizing the Beginning: Avoiding Storage Dependency for Model Merging in Continual Learning

Xi Wang, Cheng Deng

Model merging provides a compelling paradigm for integrating specialized expertise into a unified multi-task model, a goal that aligns naturally with the sequential knowledge acqui…

cs.CV2025

A Turn Toward Better Alignment: Few-Shot Generative Adaptation with Equivariant Feature Rotation

Chenghao Xu, Qi Liu, Jiexi Yan +2

Few-shot image generation aims to effectively adapt a source generative model to a target domain using very few training images. Most existing approaches introduce consistency cons…

cs.LG2024

One-Step Multi-View Clustering Based on Transition Probability

Wenhui Zhao, Quanxue Gao, Guangfei Li +2

The large-scale multi-view clustering algorithms, based on the anchor graph, have shown promising performance and efficiency and have been extensively explored in recent years. Des…

cs.CV2026

DIMoE-Adapters: Dynamic Expert Evolution for Continual Learning in Vision-Language Models

Mengxin Qin, Xiang Zhang, Xi Wang +3

Continual learning enables vision-language models to accumulate knowledge and adapt to evolving tasks without retraining from scratch. However, in multi-domain task-incremental lea…

cs.CV2020

Projection & Probability-Driven Black-Box Attack

Jie Li, Rongrong Ji, Hong Liu +4

Generating adversarial examples in a black-box setting retains a significant challenge with vast practical application prospects. In particular, existing black-box attacks suffer f…

cs.CV2019

Stacked Semantic-Guided Network for Zero-Shot Sketch-Based Image Retrieval

Hao Wang, Cheng Deng, Xinxu Xu +3

Zero-shot sketch-based image retrieval (ZS-SBIR) is a task of cross-domain image retrieval from a natural image gallery with free-hand sketch under a zero-shot scenario. Previous w…

cs.CV2024

High-Discriminative Attribute Feature Learning for Generalized Zero-Shot Learning

Yu Lei, Guoshuai Sheng, Fangfang Li +3

Zero-shot learning(ZSL) aims to recognize new classes without prior exposure to their samples, relying on semantic knowledge from observed classes. However, current attention-based…

cs.CV2019

Pyramidal Person Re-IDentification via Multi-Loss Dynamic Training

Feng Zheng, Cheng Deng, Xing Sun +5

Most existing Re-IDentification (Re-ID) methods are highly dependent on precise bounding boxes that enable images to be aligned with each other. However, due to the challenging pra…

cs.CL2025

GTA: Grouped-head latenT Attention

Luoyang Sun, Cheng Deng, Jiwen Jiang +5

Attention mechanisms underpin the success of large language models (LLMs), yet their substantial computational and memory overhead poses challenges for optimizing efficiency and pe…

cs.RO2026

-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs

Siting Wang, Xiaofeng Wang, Zheng Zhu +7

Flow-based vision-language-action (VLA) models excel in embodied control but suffer from intractable likelihoods during multi-step sampling, hindering online reinforcement learning…

cs.LG2024

Multimodal Fusion on Low-quality Data: A Comprehensive Survey

Qingyang Zhang, Yake Wei, Zongbo Han +8

Multimodal fusion focuses on integrating information from multiple modalities with the goal of more accurate prediction, which has achieved remarkable progress in a wide range of s…

cs.CV2020

Multi-task Collaborative Network for Joint Referring Expression Comprehension and Segmentation

Gen Luo, Yiyi Zhou, Xiaoshuai Sun +4

Referring expression comprehension (REC) and segmentation (RES) are two highly-related tasks, which both aim at identifying the referent according to a natural language expression.…

cs.CL2026

Learning Stateful Predictive Knowledge From Experience

Yan Song, Xidong Feng, Bo Liu +7

As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights. Viewed through the lens of predicti…

cs.CV2019

Active Transfer Learning Network: A Unified Deep Joint Spectral-Spatial Feature Learning Model For Hyperspectral Image Classification

Cheng Deng, Yumeng Xue, Xianglong Liu +2

Deep learning has recently attracted significant attention in the field of hyperspectral images (HSIs) classification. However, the construction of an efficient deep neural network…

physics.chem-ph2023

3D Molecular Geometry Analysis with 2D Graphs

Zhao Xu, Yaochen Xie, Youzhi Luo +7

Ground-state 3D geometries of molecules are essential for many molecular analysis tasks. Modern quantum mechanical methods can compute accurate 3D geometries but are computationall…

cs.CV2026

Revealing and Enhancing Core Visual Regions: Harnessing Internal Attention Dynamics for Hallucination Mitigation in LVLMs

Guangtao Lyu, Qi Liu, Chenghao Xu +5

LVLMs have achieved strong multimodal reasoning capabilities but remain prone to hallucinations, producing outputs inconsistent with visual inputs or user instructions. Existing tr…

cs.CV2019

Active Multi-Kernel Domain Adaptation for Hyperspectral Image Classification

Cheng Deng, Xianglong Liu, Chao Li +1

Recent years have witnessed the quick progress of the hyperspectral images (HSI) classification. Most of existing studies either heavily rely on the expensive label information usi…

cs.CV2019

Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning

Shaohui Lin, Rongrong Ji, Yuchao Li +2

The success of convolutional neural networks (CNNs) in computer vision applications has been accompanied by a significant increase of computation and memory costs, which prohibits…

cs.IR2019

Coupled CycleGAN: Unsupervised Hashing Network for Cross-Modal Retrieval

Chao Li, Cheng Deng, Lei Wang +2

In recent years, hashing has attracted more and more attention owing to its superior capacity of low storage cost and high query efficiency in large-scale cross-modal retrieval. Be…

cs.CV2025

Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation

Guangtao Lyu, Chenghao Xu, Qi Liu +4

Music to 3D dance generation aims to synthesize realistic and rhythmically synchronized human dance from music. While existing methods often rely on additional genre labels to furt…

cs.LG2020

Privacy-Preserving Asynchronous Federated Learning Algorithms for Multi-Party Vertically Collaborative Learning

Bin Gu, An Xu, Zhouyuan Huo +2

The privacy-preserving federated learning for vertically partitioned data has shown promising results as the solution of the emerging multi-party joint modeling application, in whi…

cs.LG2017

Deep Clustering via Joint Convolutional Autoencoder Embedding and Relative Entropy Minimization

Kamran Ghasedi Dizaji, Amirhossein Herandi, Cheng Deng +2

Image clustering is one of the most important computer vision applications, which has been extensively studied in literature. However, current clustering methods mostly suffer from…

cs.CR2022

Similarity Distribution based Membership Inference Attack on Person Re-identification

Junyao Gao, Xinyang Jiang, Huishuai Zhang +6

While person Re-identification (Re-ID) has progressed rapidly due to its wide real-world applications, it also causes severe risks of leaking personal information from training dat…

cs.CV2019

DistillHash: Unsupervised Deep Hashing by Distilling Data Pairs

Erkun Yang, Tongliang Liu, Cheng Deng +2

Due to the high storage and search efficiency, hashing has become prevalent for large-scale similarity search. Particularly, deep hashing methods have greatly improved the search p…

cs.IR2019

Query-Adaptive Hash Code Ranking for Large-Scale Multi-View Visual Search

Xianglong Liu, Lei Huang, Cheng Deng +2

Hash based nearest neighbor search has become attractive in many applications. However, the quantization in hashing usually degenerates the discriminative power when using Hamming…

cs.CV2025

A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages

Zibo Su, Kun Wei, Jiahua Li +3

Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in English but struggle with non-Engl…

cs.CV2018

Self-Supervised Adversarial Hashing Networks for Cross-Modal Retrieval

Chao Li, Cheng Deng, Ning Li +3

Thanks to the success of deep learning, cross-modal retrieval has made significant progress recently. However, there still remains a crucial bottleneck: how to bridge the modality…

cs.CV2019

Semantic Adversarial Network for Zero-Shot Sketch-Based Image Retrieval

Xinxun Xu, Hao Wang, Leida Li +1

Zero-shot sketch-based image retrieval (ZS-SBIR) is a specific cross-modal retrieval task for retrieving natural images with free-hand sketches under zero-shot scenario. Previous w…

cs.CV2025

CAMS: Towards Compositional Zero-Shot Learning via Gated Cross-Attention and Multi-Space Disentanglement

Pan Yang, Cheng Deng, Jing Yang +5

Compositional zero-shot learning (CZSL) aims to learn the concepts of attributes and objects in seen compositions and to recognize their unseen compositions. Most Contrastive Langu…

cs.CV2019

CerfGAN: A Compact, Effective, Robust, and Fast Model for Unsupervised Multi-Domain Image-to-Image Translation

Xiao Liu, Shengchuan Zhang, Hong Liu +3

In this paper, we aim at solving the multi-domain image-to-image translation problem with a unified model in an unsupervised manner. The most successful work in this area refers to…

cs.CV2025

Taming Transformer for Emotion-Controllable Talking Face Generation

Ziqi Zhang, Cheng Deng

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talki…

cs.CL2025

Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference

Qingfa Xiao, Jiachuan Wang, Haoyang Li +6

Recent advances in large language models (LLMs) have showcased exceptional performance in long-context tasks, while facing significant inference efficiency challenges with limited…

cs.IR2019

Triplet-Based Deep Hashing Network for Cross-Modal Retrieval

Cheng Deng, Zhaojia Chen, Xianglong Liu +2

Given the benefits of its low storage requirements and high retrieval efficiency, hashing has recently received increasing attention. In particular,cross-modal hashing has been wid…

cs.LG2024

Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

Muning Wen, Junwei Liao, Cheng Deng +3

Large Language Models (LLMs) have shown promise as intelligent agents in interactive decision-making tasks. Traditional approaches often depend on meticulously designed prompts, hi…

cs.LG2021

Molecule3D: A Benchmark for Predicting 3D Geometries from Molecular Graphs

Zhao Xu, Youzhi Luo, Xuan Zhang +7

Graph neural networks are emerging as promising methods for modeling molecular graphs, in which nodes and edges correspond to atoms and chemical bonds, respectively. Recent studies…

cs.LG2026

RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis

Zhen Bi, Xueshu Chen, Luoyang Sun +4

The transition toward localized intelligence through Small Language Models (SLMs) has intensified the need for rigorous performance characterization on resource-constrained edge ha…

cs.CV2026

VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models

Huawei Ji, Yuanhao Sun, Yuan Jin +4

Visual token pruning methods effectively mitigate the quadratic computational growth caused by processing high-resolution images and video frames in vision-language models (VLMs).…

cs.CV2021

Doubly Contrastive Deep Clustering

Zhiyuan Dang, Cheng Deng, Xu Yang +1

Deep clustering successfully provides more effective features than conventional ones and thus becomes an important technique in current unsupervised learning. However, most deep cl…

cs.CV2025

Knowledge-Guided Prompt Learning for Deepfake Facial Image Detection

Hao Wang, Cheng Deng, Zhidong Zhao

Recent generative models demonstrate impressive performance on synthesizing photographic images, which makes humans hardly to distinguish them from pristine ones, especially on rea…

cs.AI2026

Logical Structure as Knowledge: Enhancing LLM Reasoning via Structured Logical Knowledge Density Estimation

Zhen Bi, Zhenlin Hu, Xueshu Chen +7

The reasoning capabilities of Large Language Models (LLMs) are increasingly attributed to training data quality rather than mere parameter scaling. However, existing data-centric p…